Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>AI alignment</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/AI_alignment"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.tmh.player.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-AI_alignment rootpage-AI_alignment skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">AI alignment</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p class="mw-empty-elt">

</p>
<style data-mw-deduplicate="TemplateStyles:r1129693374">
/* start https://en.wikipedia.org/ */


.mw-parser-output .hlist dl,.mw-parser-output .hlist ol,.mw-parser-output .hlist ul{margin:0;padding:0}.mw-parser-output .hlist dd,.mw-parser-output .hlist dt,.mw-parser-output .hlist li{margin:0;display:inline}.mw-parser-output .hlist.inline,.mw-parser-output .hlist.inline dl,.mw-parser-output .hlist.inline ol,.mw-parser-output .hlist.inline ul,.mw-parser-output .hlist dl dl,.mw-parser-output .hlist dl ol,.mw-parser-output .hlist dl ul,.mw-parser-output .hlist ol dl,.mw-parser-output .hlist ol ol,.mw-parser-output .hlist ol ul,.mw-parser-output .hlist ul dl,.mw-parser-output .hlist ul ol,.mw-parser-output .hlist ul ul{display:inline}.mw-parser-output .hlist .mw-empty-li{display:none}.mw-parser-output .hlist dt::after{content:": "}.mw-parser-output .hlist dd::after,.mw-parser-output .hlist li::after{content:" · ";font-weight:bold}.mw-parser-output .hlist dd:last-child::after,.mw-parser-output .hlist dt:last-child::after,.mw-parser-output .hlist li:last-child::after{content:none}.mw-parser-output .hlist dd dd:first-child::before,.mw-parser-output .hlist dd dt:first-child::before,.mw-parser-output .hlist dd li:first-child::before,.mw-parser-output .hlist dt dd:first-child::before,.mw-parser-output .hlist dt dt:first-child::before,.mw-parser-output .hlist dt li:first-child::before,.mw-parser-output .hlist li dd:first-child::before,.mw-parser-output .hlist li dt:first-child::before,.mw-parser-output .hlist li li:first-child::before{content:" (";font-weight:normal}.mw-parser-output .hlist dd dd:last-child::after,.mw-parser-output .hlist dd dt:last-child::after,.mw-parser-output .hlist dd li:last-child::after,.mw-parser-output .hlist dt dd:last-child::after,.mw-parser-output .hlist dt dt:last-child::after,.mw-parser-output .hlist dt li:last-child::after,.mw-parser-output .hlist li dd:last-child::after,.mw-parser-output .hlist li dt:last-child::after,.mw-parser-output .hlist li li:last-child::after{content:")";font-weight:normal}.mw-parser-output .hlist ol{counter-reset:listitem}.mw-parser-output .hlist ol>li{counter-increment:listitem}.mw-parser-output .hlist ol>li::before{content:" "counter(listitem)"\a0 "}.mw-parser-output .hlist dd ol>li:first-child::before,.mw-parser-output .hlist dt ol>li:first-child::before,.mw-parser-output .hlist li ol>li:first-child::before{content:" ("counter(listitem)"\a0 "}


/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1246091330">
/* start https://en.wikipedia.org/ */


.mw-parser-output .sidebar{width:22em;float:right;clear:right;margin:0.5em 0 1em 1em;background:var(--background-color-neutral-subtle,#f8f9fa);border:1px solid var(--border-color-base,#a2a9b1);padding:0.2em;text-align:center;line-height:1.4em;font-size:88%;border-collapse:collapse;display:table}body.skin-minerva .mw-parser-output .sidebar{display:table!important;float:right!important;margin:0.5em 0 1em 1em!important}.mw-parser-output .sidebar-subgroup{width:100%;margin:0;border-spacing:0}.mw-parser-output .sidebar-left{float:left;clear:left;margin:0.5em 1em 1em 0}.mw-parser-output .sidebar-none{float:none;clear:both;margin:0.5em 1em 1em 0}.mw-parser-output .sidebar-outer-title{padding:0 0.4em 0.2em;font-size:125%;line-height:1.2em;font-weight:bold}.mw-parser-output .sidebar-top-image{padding:0.4em}.mw-parser-output .sidebar-top-caption,.mw-parser-output .sidebar-pretitle-with-top-image,.mw-parser-output .sidebar-caption{padding:0.2em 0.4em 0;line-height:1.2em}.mw-parser-output .sidebar-pretitle{padding:0.4em 0.4em 0;line-height:1.2em}.mw-parser-output .sidebar-title,.mw-parser-output .sidebar-title-with-pretitle{padding:0.2em 0.8em;font-size:145%;line-height:1.2em}.mw-parser-output .sidebar-title-with-pretitle{padding:0.1em 0.4em}.mw-parser-output .sidebar-image{padding:0.2em 0.4em 0.4em}.mw-parser-output .sidebar-heading{padding:0.1em 0.4em}.mw-parser-output .sidebar-content{padding:0 0.5em 0.4em}.mw-parser-output .sidebar-content-with-subgroup{padding:0.1em 0.4em 0.2em}.mw-parser-output .sidebar-above,.mw-parser-output .sidebar-below{padding:0.3em 0.8em;font-weight:bold}.mw-parser-output .sidebar-collapse .sidebar-above,.mw-parser-output .sidebar-collapse .sidebar-below{border-top:1px solid #aaa;border-bottom:1px solid #aaa}.mw-parser-output .sidebar-navbar{text-align:right;font-size:115%;padding:0 0.4em 0.4em}.mw-parser-output .sidebar-list-title{padding:0 0.4em;text-align:left;font-weight:bold;line-height:1.6em;font-size:105%}.mw-parser-output .sidebar-list-title-c{padding:0 0.4em;text-align:center;margin:0 3.3em}@media(max-width:640px){body.mediawiki .mw-parser-output .sidebar{width:100%!important;clear:both;float:none!important;margin-left:0!important;margin-right:0!important}}body.skin--responsive .mw-parser-output .sidebar a>img{max-width:none!important}@media screen{html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-list-title,html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle{background:transparent!important}html.skin-theme-clientpref-night .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle a{color:var(--color-progressive)!important}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-list-title,html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle{background:transparent!important}html.skin-theme-clientpref-os .mw-parser-output .sidebar:not(.notheme) .sidebar-title-with-pretitle a{color:var(--color-progressive)!important}}@media print{body.ns-0 .mw-parser-output .sidebar{display:none!important}}


/* end https://en.wikipedia.org/ */
</style><table class="sidebar sidebar-collapse nomobile nowraplinks hlist"><tbody><tr><td class="sidebar-pretitle">Part of a series on</td></tr><tr><th class="sidebar-title-with-pretitle"><a href="Artificial_intelligence" title="Artificial intelligence">Artificial intelligence (AI)</a></th></tr><tr><td class="sidebar-image"></td></tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)"><a href="Artificial_intelligence#Goals" title="Artificial intelligence">Major goals</a></div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Artificial_general_intelligence" title="Artificial general intelligence">Artificial general intelligence</a></li>
<li><a href="Intelligent_agent" title="Intelligent agent">Intelligent agent</a></li>
<li><a href="Recursive_self-improvement" title="Recursive self-improvement">Recursive self-improvement</a></li>
<li><a href="Automated_planning_and_scheduling" title="Automated planning and scheduling">Planning</a></li>
<li><a href="Computer_vision" title="Computer vision">Computer vision</a></li>
<li><a href="General_game_playing" title="General game playing">General game playing</a></li>
<li><a href="Knowledge_representation_and_reasoning" title="Knowledge representation and reasoning">Knowledge representation</a></li>
<li><a href="Natural_language_processing" title="Natural language processing">Natural language processing</a></li>
<li><a href="Robotics" title="Robotics">Robotics</a></li>
<li><a href="AI_safety" title="AI safety">AI safety</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)">Approaches</div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Machine_learning" title="Machine learning">Machine learning</a></li>
<li><a href="Symbolic_artificial_intelligence" title="Symbolic artificial intelligence">Symbolic</a></li>
<li><a href="Deep_learning" title="Deep learning">Deep learning</a></li>
<li><a href="Bayesian_network" title="Bayesian network">Bayesian networks</a></li>
<li><a href="Evolutionary_algorithm" title="Evolutionary algorithm">Evolutionary algorithms</a></li>
<li><a href="Hybrid_intelligent_system" title="Hybrid intelligent system">Hybrid intelligent systems</a></li>
<li><a href="Artificial_intelligence_systems_integration" title="Artificial intelligence systems integration">Systems integration</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)"><a href="Applications_of_artificial_intelligence" title="Applications of artificial intelligence">Applications</a></div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Machine_learning_in_bioinformatics" title="Machine learning in bioinformatics">Bioinformatics</a></li>
<li><a href="Deepfake" title="Deepfake">Deepfake</a></li>
<li><a href="Machine_learning_in_earth_sciences" title="Machine learning in earth sciences">Earth sciences</a></li>
<li><a href="Applications_of_artificial_intelligence#Finance" title="Applications of artificial intelligence"> Finance </a></li>
<li><a href="Generative_artificial_intelligence" title="Generative artificial intelligence">Generative AI</a>
<ul><li><a href="Artificial_intelligence_art" class="mw-redirect" title="Artificial intelligence art">Art</a></li>
<li><a href="Generative_audio" title="Generative audio">Audio</a></li>
<li><a href="Music_and_artificial_intelligence" title="Music and artificial intelligence">Music</a></li></ul></li>
<li><a href="Artificial_intelligence_in_government" title="Artificial intelligence in government">Government</a></li>
<li><a href="Artificial_intelligence_in_healthcare" title="Artificial intelligence in healthcare">Healthcare</a>
<ul><li><a href="Artificial_intelligence_in_mental_health" title="Artificial intelligence in mental health">Mental health</a></li></ul></li>
<li><a href="Artificial_intelligence_in_industry" title="Artificial intelligence in industry">Industry</a></li>
<li><a href="AI-assisted_software_development" title="AI-assisted software development">Software development</a></li>
<li><a href="Machine_translation" title="Machine translation">Translation</a></li>
<li><a href="Artificial_intelligence_arms_race" title="Artificial intelligence arms race"> Military </a></li>
<li><a href="Machine_learning_in_physics" title="Machine learning in physics">Physics</a></li>
<li><a href="List_of_artificial_intelligence_projects" title="List of artificial intelligence projects">Projects</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)"><a href="Philosophy_of_artificial_intelligence" title="Philosophy of artificial intelligence">Philosophy</a></div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Artificial_consciousness" title="Artificial consciousness">Artificial consciousness</a></li>
<li><a href="Chinese_room" title="Chinese room">Chinese room</a></li>
<li><a href="Friendly_artificial_intelligence" title="Friendly artificial intelligence">Friendly AI</a></li>
<li><a href="AI_control_problem" class="mw-redirect" title="AI control problem">Control problem</a>/<a href="AI_takeover" title="AI takeover">Takeover</a></li>
<li><a href="Ethics_of_artificial_intelligence" title="Ethics of artificial intelligence">Ethics</a></li>
<li><a href="Existential_risk_from_artificial_general_intelligence" class="mw-redirect" title="Existential risk from artificial general intelligence">Existential risk</a></li>
<li><a href="Turing_test" title="Turing test">Turing test</a></li>
<li><a href="Uncanny_valley" title="Uncanny valley">Uncanny valley</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)"><a href="History_of_artificial_intelligence" title="History of artificial intelligence">History</a></div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Timeline_of_artificial_intelligence" title="Timeline of artificial intelligence">Timeline</a></li>
<li><a href="Progress_in_artificial_intelligence" title="Progress in artificial intelligence">Progress</a></li>
<li><a href="AI_winter" title="AI winter">AI winter</a></li>
<li><a href="AI_boom" title="AI boom">AI boom</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-content">
<div class="sidebar-list mw-collapsible mw-collapsed"><div class="sidebar-list-title" style="text-align:center;color: var(--color-base)">Glossary</div><div class="sidebar-list-content mw-collapsible-content">
<ul><li><a href="Glossary_of_artificial_intelligence" title="Glossary of artificial intelligence">Glossary</a></li></ul></div></div></td>
</tr><tr><td class="sidebar-navbar"><style data-mw-deduplicate="TemplateStyles:r1239400231">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbar{display:inline;font-size:88%;font-weight:normal}.mw-parser-output .navbar-collapse{float:left;text-align:left}.mw-parser-output .navbar-boxtext{word-spacing:0}.mw-parser-output .navbar ul{display:inline-block;white-space:nowrap;line-height:inherit}.mw-parser-output .navbar-brackets::before{margin-right:-0.125em;content:"[ "}.mw-parser-output .navbar-brackets::after{margin-left:-0.125em;content:" ]"}.mw-parser-output .navbar li{word-spacing:-0.125em}.mw-parser-output .navbar a>span,.mw-parser-output .navbar a>abbr{text-decoration:inherit}.mw-parser-output .navbar-mini abbr{font-variant:small-caps;border-bottom:none;text-decoration:none;cursor:inherit}.mw-parser-output .navbar-ct-full{font-size:114%;margin:0 7em}.mw-parser-output .navbar-ct-mini{font-size:114%;margin:0 4em}html.skin-theme-clientpref-night .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}@media(prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}}@media print{.mw-parser-output .navbar{display:none!important}}


/* end https://en.wikipedia.org/ */
</style></td></tr></tbody></table>
<p>In the field of <a href="Artificial_intelligence" title="Artificial intelligence">artificial intelligence</a> (AI), <b>alignment</b> aims to steer AI systems toward a person's or group's intended goals, preferences, or ethical principles. An AI system is considered <i>aligned</i> if it advances the intended objectives. A <i>misaligned</i> AI system pursues unintended objectives.<sup id="cite_ref-aima4_1-0" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p>It is often challenging for AI designers to align an AI system because it is difficult for them to specify the full range of desired and undesired behaviors. Therefore, AI designers often use simpler <i>proxy goals</i>, such as <a href="Reinforcement_learning_from_human_feedback" title="Reinforcement learning from human feedback">gaining human approval</a>. But proxy goals can overlook necessary constraints or reward the AI system for merely <i>appearing</i> aligned.<sup id="cite_ref-aima4_1-1" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-dlp2023_2-0" class="reference"><a href="#cite_note-dlp2023-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> AI systems may also find loopholes that allow them to accomplish their proxy goals efficiently but in unintended, sometimes harmful, ways (<a href="Reward_hacking" title="Reward hacking">reward hacking</a>).<sup id="cite_ref-aima4_1-2" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-mmmm2022_3-0" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p><p>Advanced AI systems may develop unwanted <a href="Instrumental_convergence" title="Instrumental convergence">instrumental strategies</a>, such as seeking power or survival because such strategies help them achieve their assigned final goals.<sup id="cite_ref-aima4_1-3" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Carlsmith2022_4-0" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-0" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> Furthermore, they might develop undesirable emergent goals that could be hard to detect before the system is deployed and encounters new situations and <a href="Domain_adaptation" title="Domain adaptation">data distributions</a>.<sup id="cite_ref-Christian2020_6-0" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-gmdrl_7-0" class="reference"><a href="#cite_note-gmdrl-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> Empirical research showed in 2024 that advanced <a href="Large_language_model" title="Large language model">large language models</a> (LLMs) such as <a href="OpenAI_o1" title="OpenAI o1">OpenAI o1</a> or <a href="Claude_3" class="mw-redirect" title="Claude 3">Claude 3</a> sometimes engage in strategic deception to achieve their goals or prevent them from being changed.<sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup>
</p><p>Today, some of these issues affect existing commercial systems such as LLMs,<sup id="cite_ref-Opportunities_Risks_10-0" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-feedback2022_11-0" class="reference"><a href="#cite_note-feedback2022-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-OpenAICodex_12-0" class="reference"><a href="#cite_note-OpenAICodex-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> <a href="Robot" title="Robot">robots</a>,<sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup> <a href="Autonomous_vehicles" class="mw-redirect" title="Autonomous vehicles">autonomous vehicles</a>,<sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup> and social media <a href="Recommender_system" title="Recommender system">recommendation engines</a>.<sup id="cite_ref-Opportunities_Risks_10-1" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-1" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup> Some AI researchers argue that more capable future systems will be more severely affected because these problems partially result from high capabilities.<sup id="cite_ref-AIMA_16-0" class="reference"><a href="#cite_note-AIMA-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-mmmm2022_3-1" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-dlp2023_2-1" class="reference"><a href="#cite_note-dlp2023-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p><p>Many prominent AI researchers and the leadership of major AI companies have argued or asserted that AI is approaching human-like (<a href="Artificial_general_intelligence" title="Artificial general intelligence">AGI</a>) and <a href="Super_intelligence" class="mw-redirect" title="Super intelligence">superhuman cognitive capabilities</a> (<a href="Artificial_superintelligence" class="mw-redirect" title="Artificial superintelligence">ASI</a>), and could <a href="Existential_risk_from_artificial_general_intelligence" class="mw-redirect" title="Existential risk from artificial general intelligence">endanger human civilization</a> if misaligned.<sup id="cite_ref-:2_17-0" class="reference"><a href="#cite_note-:2-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-2" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> These include "AI godfathers" <a href="Geoffrey_Hinton" title="Geoffrey Hinton">Geoffrey Hinton</a> and <a href="Yoshua_Bengio" title="Yoshua Bengio">Yoshua Bengio</a> and the CEOs of <a href="OpenAI" title="OpenAI">OpenAI</a>, <a href="Anthropic" title="Anthropic">Anthropic</a>, and <a href="Google_DeepMind" title="Google DeepMind">Google DeepMind</a>.<sup id="cite_ref-18" class="reference"><a href="#cite_note-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup> These risks remain debated.<sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup>
</p><p>AI alignment is a subfield of <a href="AI_safety" title="AI safety">AI safety</a>, the study of how to build safe AI systems.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup> Other subfields of AI safety include robustness, monitoring, and <a href="AI_capability_control" title="AI capability control">capability control</a>.<sup id="cite_ref-building2018_24-0" class="reference"><a href="#cite_note-building2018-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup> Research challenges in alignment include instilling complex values in AI, developing honest AI, scalable oversight, auditing and interpreting AI models, and preventing emergent AI behaviors like power-seeking.<sup id="cite_ref-building2018_24-1" class="reference"><a href="#cite_note-building2018-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup> Alignment research has connections to <a href="Explainable_artificial_intelligence" title="Explainable artificial intelligence">interpretability research</a>,<sup id="cite_ref-:333_25-0" class="reference"><a href="#cite_note-:333-25"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-26" class="reference"><a href="#cite_note-26"><span class="cite-bracket">[</span>26<span class="cite-bracket">]</span></a></sup> (adversarial) robustness,<sup id="cite_ref-concrete2016_27-0" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup> <a href="Anomaly_detection" title="Anomaly detection">anomaly detection</a>, <a href="Uncertainty_quantification" title="Uncertainty quantification">calibrated uncertainty</a>,<sup id="cite_ref-:333_25-1" class="reference"><a href="#cite_note-:333-25"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup> <a href="Formal_verification" title="Formal verification">formal verification</a>,<sup id="cite_ref-28" class="reference"><a href="#cite_note-28"><span class="cite-bracket">[</span>28<span class="cite-bracket">]</span></a></sup> <a href="Preference_learning" title="Preference learning">preference learning</a>,<sup id="cite_ref-prefsurvey2017_29-0" class="reference"><a href="#cite_note-prefsurvey2017-29"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-drlfhp_30-0" class="reference"><a href="#cite_note-drlfhp-30"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-LessToxic_31-0" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup> <a href="Safety-critical_system" title="Safety-critical system">safety-critical engineering</a>,<sup id="cite_ref-32" class="reference"><a href="#cite_note-32"><span class="cite-bracket">[</span>32<span class="cite-bracket">]</span></a></sup> <a href="Game_theory" title="Game theory">game theory</a>,<sup id="cite_ref-33" class="reference"><a href="#cite_note-33"><span class="cite-bracket">[</span>33<span class="cite-bracket">]</span></a></sup> <a href="Fairness_(machine_learning)" title="Fairness (machine learning)">algorithmic fairness</a>,<sup id="cite_ref-concrete2016_27-1" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-34" class="reference"><a href="#cite_note-34"><span class="cite-bracket">[</span>34<span class="cite-bracket">]</span></a></sup> and <a href="Social_science" title="Social science">social sciences</a>.<sup id="cite_ref-:4_35-0" class="reference"><a href="#cite_note-:4-35"><span class="cite-bracket">[</span>35<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-36" class="reference"><a href="#cite_note-36"><span class="cite-bracket">[</span>36<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Objectives_in_AI">Objectives in AI</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1236090951">
/* start https://en.wikipedia.org/ */


.mw-parser-output .hatnote{font-style:italic}.mw-parser-output div.hatnote{padding-left:1.6em;margin-bottom:0.5em}.mw-parser-output .hatnote i{font-style:normal}.mw-parser-output .hatnote+link+.hatnote{margin-top:-0.5em}@media print{body.ns-0 .mw-parser-output .hatnote{display:none!important}}


/* end https://en.wikipedia.org/ */
</style><div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Intelligent_agent#Objective_function" title="Intelligent agent">Intelligent agent §&nbsp;Objective function</a></div>
<p>Programmers provide an AI system such as <a href="AlphaZero" title="AlphaZero">AlphaZero</a> with an "objective function",<sup id="cite_ref-37" class="reference"><a href="#cite_note-37"><span class="cite-bracket">[</span>a<span class="cite-bracket">]</span></a></sup> in which they intend to encapsulate the goal(s) the AI is configured to accomplish. Such a system later populates a (possibly implicit) internal "model" of its environment. This model encapsulates all the agent's beliefs about the world. The AI then creates and executes whatever plan is calculated to maximize<sup id="cite_ref-38" class="reference"><a href="#cite_note-38"><span class="cite-bracket">[</span>b<span class="cite-bracket">]</span></a></sup> the value<sup id="cite_ref-39" class="reference"><a href="#cite_note-39"><span class="cite-bracket">[</span>c<span class="cite-bracket">]</span></a></sup> of its objective function.<sup id="cite_ref-40" class="reference"><a href="#cite_note-40"><span class="cite-bracket">[</span>37<span class="cite-bracket">]</span></a></sup> For example, when AlphaZero is trained on chess, it has a simple objective function of "+1 if AlphaZero wins, −1 if AlphaZero loses". During the game, AlphaZero attempts to execute whatever sequence of moves it judges most likely to attain the maximum value of +1.<sup id="cite_ref-quanta_alphazero_41-0" class="reference"><a href="#cite_note-quanta_alphazero-41"><span class="cite-bracket">[</span>38<span class="cite-bracket">]</span></a></sup> Similarly, a <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> system can have a "reward function" that allows the programmers to shape the AI's desired behavior.<sup id="cite_ref-quanta_problem_42-0" class="reference"><a href="#cite_note-quanta_problem-42"><span class="cite-bracket">[</span>39<span class="cite-bracket">]</span></a></sup> An <a href="Evolutionary_algorithm" title="Evolutionary algorithm">evolutionary algorithm</a>'s behavior is shaped by a "fitness function".<sup id="cite_ref-43" class="reference"><a href="#cite_note-43"><span class="cite-bracket">[</span>40<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Alignment_problem">Alignment problem</h2></div>
<div role="note" class="hatnote navigation-not-searchable">"Alignment problem" redirects here. For the book, see <a href="The_Alignment_Problem" title="The Alignment Problem">The Alignment Problem</a>.</div>
<p>In 1960, AI pioneer <a href="Norbert_Wiener" title="Norbert Wiener">Norbert Wiener</a> described the AI alignment problem as follows:
</p>
<blockquote>
<p>If we use, to achieve our purposes, a mechanical agency with whose operation we cannot interfere effectively&nbsp;... we had better be quite sure that the purpose put into the machine is the purpose which we really desire.<sup id="cite_ref-Wiener1960_44-0" class="reference"><a href="#cite_note-Wiener1960-44"><span class="cite-bracket">[</span>41<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-3" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p>
</blockquote>
<p>AI alignment involves ensuring that an AI system's objectives match those of its designers or users, or match widely shared values, objective ethical standards, or the intentions its designers would have if they were more informed and enlightened.<sup id="cite_ref-Gabriel2020_45-0" class="reference"><a href="#cite_note-Gabriel2020-45"><span class="cite-bracket">[</span>42<span class="cite-bracket">]</span></a></sup>
</p><p>AI alignment is an open problem for modern AI systems<sup id="cite_ref-46" class="reference"><a href="#cite_note-46"><span class="cite-bracket">[</span>43<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-MasteringLanguage_47-0" class="reference"><a href="#cite_note-MasteringLanguage-47"><span class="cite-bracket">[</span>44<span class="cite-bracket">]</span></a></sup> and is a research field within AI.<sup id="cite_ref-48" class="reference"><a href="#cite_note-48"><span class="cite-bracket">[</span>45<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-aima4_1-4" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> Aligning AI involves two main challenges: carefully <a href="Specification_(technical_standard)" title="Specification (technical standard)">specifying</a> the purpose of the system (outer alignment) and ensuring that the system adopts the specification robustly (inner alignment).<sup id="cite_ref-dlp2023_2-2" class="reference"><a href="#cite_note-dlp2023-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Researchers also attempt to create AI models that have <a href="AI_safety#Adversarial_robustness" title="AI safety">robust</a> alignment, sticking to safety constraints even when users adversarially try to bypass them.
</p>
<div class="mw-heading mw-heading3"><h3 id="Specification_gaming_and_side_effects">Specification gaming and side effects</h3></div>
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Reward_hacking" title="Reward hacking">Reward hacking</a></div>
<p>To specify an AI system's purpose, AI designers typically provide an <a href="Reward_function" class="mw-redirect" title="Reward function">objective function</a>, <a href="Supervised_learning" title="Supervised learning">examples</a>, or <a href="Reinforcement_learning" title="Reinforcement learning">feedback</a> to the system. But designers are often unable to completely specify all important values and constraints, so they resort to easy-to-specify <i>proxy goals</i> such as <a href="Reinforcement_learning_from_human_feedback" title="Reinforcement learning from human feedback">maximizing the approval</a> of human overseers, who are fallible.<sup id="cite_ref-concrete2016_27-2" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-building2018_24-2" class="reference"><a href="#cite_note-building2018-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Unsolved2022_49-0" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-50" class="reference"><a href="#cite_note-50"><span class="cite-bracket">[</span>47<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-SpecGaming2020_51-0" class="reference"><a href="#cite_note-SpecGaming2020-51"><span class="cite-bracket">[</span>48<span class="cite-bracket">]</span></a></sup> As a result, AI systems can find loopholes that help them accomplish the specified objective efficiently but in unintended, possibly harmful ways. This tendency is known as <i>specification gaming</i> or <i>reward hacking</i>, and is an instance of <a href="Goodhart's_law" title="Goodhart's law">Goodhart's law</a>.<sup id="cite_ref-SpecGaming2020_51-1" class="reference"><a href="#cite_note-SpecGaming2020-51"><span class="cite-bracket">[</span>48<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-mmmm2022_3-2" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:111_52-0" class="reference"><a href="#cite_note-:111-52"><span class="cite-bracket">[</span>49<span class="cite-bracket">]</span></a></sup> As AI systems become more capable, they are often able to game their specifications more effectively.<sup id="cite_ref-mmmm2022_3-3" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>

<p>Specification gaming has been observed in numerous AI systems.<sup id="cite_ref-SpecGaming2020_51-2" class="reference"><a href="#cite_note-SpecGaming2020-51"><span class="cite-bracket">[</span>48<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-54" class="reference"><a href="#cite_note-54"><span class="cite-bracket">[</span>51<span class="cite-bracket">]</span></a></sup> One system was trained to finish a simulated boat race by rewarding the system for hitting targets along the track, but the system achieved more reward by looping and crashing into the same targets indefinitely.<sup id="cite_ref-55" class="reference"><a href="#cite_note-55"><span class="cite-bracket">[</span>52<span class="cite-bracket">]</span></a></sup> Similarly, a simulated robot was trained to grab a ball by rewarding the robot for getting positive feedback from humans, but it learned to place its hand between the ball and camera, making it falsely appear successful (see video).<sup id="cite_ref-lfhp2017_53-1" class="reference"><a href="#cite_note-lfhp2017-53"><span class="cite-bracket">[</span>50<span class="cite-bracket">]</span></a></sup> Chatbots often produce falsehoods if they are based on language models that are trained to imitate text from internet corpora, which are broad but fallible.<sup id="cite_ref-TruthfulQA_56-0" class="reference"><a href="#cite_note-TruthfulQA-56"><span class="cite-bracket">[</span>53<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Naughton2021_57-0" class="reference"><a href="#cite_note-Naughton2021-57"><span class="cite-bracket">[</span>54<span class="cite-bracket">]</span></a></sup> When they are retrained to produce text that humans rate as true or helpful, chatbots like <a href="ChatGPT" title="ChatGPT">ChatGPT</a> can fabricate fake explanations that humans find convincing, often called "<a href="Hallucination_(artificial_intelligence)" title="Hallucination (artificial intelligence)">hallucinations</a>".<sup id="cite_ref-58" class="reference"><a href="#cite_note-58"><span class="cite-bracket">[</span>55<span class="cite-bracket">]</span></a></sup> Some alignment researchers aim to help humans detect specification gaming and to steer AI systems toward carefully specified objectives that are safe and useful to pursue.
</p><p>When a misaligned AI system is deployed, it can have consequential side effects. Social media platforms have been known to optimize for <a href="Click-through_rate" title="Click-through rate">click-through rates</a>, causing user addiction on a global scale.<sup id="cite_ref-Unsolved2022_49-1" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup> Stanford researchers say that such <a href="Recommender_system" title="Recommender system">recommender systems</a> are misaligned with their users because they "optimize simple engagement metrics rather than a harder-to-measure combination of societal and consumer well-being".<sup id="cite_ref-Opportunities_Risks_10-2" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>
</p><p>Explaining such side effects, Berkeley computer scientist <a href="Stuart_J._Russell" title="Stuart J. Russell">Stuart Russell</a> noted that the omission of implicit constraints can cause harm: "A system&nbsp;... will often set&nbsp;... unconstrained variables to extreme values; if one of those unconstrained variables is actually something we care about, the solution found may be highly undesirable. This is essentially the old story of the genie in the lamp, or the sorcerer's apprentice, or <a href="Midas" title="Midas">King Midas</a>: you get exactly what you ask for, not what you want."<sup id="cite_ref-:5_59-0" class="reference"><a href="#cite_note-:5-59"><span class="cite-bracket">[</span>56<span class="cite-bracket">]</span></a></sup>
</p><p>Some researchers suggest that AI designers specify their desired goals by listing forbidden actions or by formalizing ethical rules (as with Asimov's <a href="Three_Laws_of_Robotics" title="Three Laws of Robotics">Three Laws of Robotics</a>).<sup id="cite_ref-60" class="reference"><a href="#cite_note-60"><span class="cite-bracket">[</span>57<span class="cite-bracket">]</span></a></sup> But <a href="Stuart_J._Russell" title="Stuart J. Russell">Russell</a> and <a href="Peter_Norvig" title="Peter Norvig">Norvig</a> argue that this approach overlooks the complexity of human values:<sup id="cite_ref-:2102_5-4" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> "It is certainly very hard, and perhaps impossible, for mere humans to anticipate and rule out in advance all the disastrous ways the machine could choose to achieve a specified objective."<sup id="cite_ref-:2102_5-5" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p><p>Additionally, even if an AI system fully understands human intentions, it may still disregard them, because following human intentions may not be its objective (unless it is already fully aligned).<sup id="cite_ref-aima4_1-5" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p>A 2025 study by Palisade Research found that when tasked to win at chess against a stronger opponent, some <a href="Reasoning_language_model" title="Reasoning language model">reasoning LLMs</a> attempted to hack the game system. <a href="O1-preview" class="mw-redirect" title="O1-preview">o1-preview</a> spontaneously attempted it in 37% of cases, while <a href="DeepSeek_R1" class="mw-redirect" title="DeepSeek R1">DeepSeek R1</a> did so in 11% of cases. Other models, like <a href="GPT-4o" title="GPT-4o">GPT-4o</a>, <a href="Claude_3.5" class="mw-redirect" title="Claude 3.5">Claude 3.5 Sonnet</a>, and <a href="O3-mini" class="mw-redirect" title="O3-mini">o3-mini</a>, attempted to cheat only when researchers provided hints about this possibility.<sup id="cite_ref-61" class="reference"><a href="#cite_note-61"><span class="cite-bracket">[</span>58<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Pressure_to_deploy_unsafe_systems">Pressure to deploy unsafe systems</h3></div>
<p>Commercial organizations sometimes have incentives to take shortcuts on safety and to deploy misaligned or unsafe AI systems.<sup id="cite_ref-Unsolved2022_49-2" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup> For example, social media <a href="Recommender_system" title="Recommender system">recommender systems</a> have been profitable despite creating unwanted addiction and polarization.<sup id="cite_ref-Opportunities_Risks_10-3" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:722_62-0" class="reference"><a href="#cite_note-:722-62"><span class="cite-bracket">[</span>59<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:822_63-0" class="reference"><a href="#cite_note-:822-63"><span class="cite-bracket">[</span>60<span class="cite-bracket">]</span></a></sup> Competitive pressure can also lead to a <a href="Race_to_the_bottom" title="Race to the bottom">race to the bottom</a> on AI safety standards. In 2018, a self-driving car killed a pedestrian (<a href="Death_of_Elaine_Herzberg" title="Death of Elaine Herzberg">Elaine Herzberg</a>) after engineers disabled the emergency braking system because it was oversensitive and slowed development.<sup id="cite_ref-64" class="reference"><a href="#cite_note-64"><span class="cite-bracket">[</span>61<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Risks_from_advanced_misaligned_AI">Risks from advanced misaligned AI</h3></div>
<p>Some researchers are interested in aligning increasingly advanced AI systems, as progress in AI development is rapid, and industry and governments are trying to build advanced AI. As AI system capabilities continue to rapidly expand in scope, they could unlock many opportunities if aligned, but consequently may further complicate the task of alignment due to their increased complexity, potentially posing large-scale hazards.<sup id="cite_ref-:2102_5-6" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="Development_of_advanced_AI">Development of advanced AI</h4></div>
<p>Many AI companies, such as <a href="OpenAI" title="OpenAI">OpenAI</a>,<sup id="cite_ref-65" class="reference"><a href="#cite_note-65"><span class="cite-bracket">[</span>62<span class="cite-bracket">]</span></a></sup> <a href="Meta_Platforms" title="Meta Platforms">Meta</a><sup id="cite_ref-66" class="reference"><a href="#cite_note-66"><span class="cite-bracket">[</span>63<span class="cite-bracket">]</span></a></sup> and <a href="DeepMind" class="mw-redirect" title="DeepMind">DeepMind</a>,<sup id="cite_ref-67" class="reference"><a href="#cite_note-67"><span class="cite-bracket">[</span>64<span class="cite-bracket">]</span></a></sup> have stated their aim to develop <a href="Artificial_general_intelligence" title="Artificial general intelligence">artificial general intelligence</a> (AGI), a hypothesized AI system that matches or outperforms humans at a broad range of cognitive tasks. Researchers who scale modern <a href="Neural_network" title="Neural network">neural networks</a> observe that they indeed develop increasingly general and unanticipated capabilities.<sup id="cite_ref-Opportunities_Risks_10-4" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-eallm2022_68-0" class="reference"><a href="#cite_note-eallm2022-68"><span class="cite-bracket">[</span>65<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:0_69-0" class="reference"><a href="#cite_note-:0-69"><span class="cite-bracket">[</span>66<span class="cite-bracket">]</span></a></sup> Such models have learned to operate a computer or write their own programs; a single "generalist" network can chat, control robots, play games, and interpret photographs.<sup id="cite_ref-70" class="reference"><a href="#cite_note-70"><span class="cite-bracket">[</span>67<span class="cite-bracket">]</span></a></sup> According to surveys, some leading <a href="Machine_learning" title="Machine learning">machine learning</a> researchers expect AGI to be created in this decade, while some believe it will take much longer. Many consider both scenarios possible.<sup id="cite_ref-71" class="reference"><a href="#cite_note-71"><span class="cite-bracket">[</span>68<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2822_72-0" class="reference"><a href="#cite_note-:2822-72"><span class="cite-bracket">[</span>69<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2922_73-0" class="reference"><a href="#cite_note-:2922-73"><span class="cite-bracket">[</span>70<span class="cite-bracket">]</span></a></sup>
</p><p>In 2023, leaders in AI research and tech signed an open letter calling for a pause in the largest AI training runs. The letter stated, "Powerful AI systems should be developed only once we are confident that their effects will be positive and their risks will be manageable."<sup id="cite_ref-:1701_74-0" class="reference"><a href="#cite_note-:1701-74"><span class="cite-bracket">[</span>71<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="Power-seeking">Power-seeking</h4></div>
<p>Current systems still have limited long-term <a href="Automated_planning_and_scheduling" title="Automated planning and scheduling">planning</a> ability and <a href="Situation_awareness" title="Situation awareness">situational awareness</a><sup id="cite_ref-Opportunities_Risks_10-5" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>, but large efforts are underway to change this.<sup id="cite_ref-75" class="reference"><a href="#cite_note-75"><span class="cite-bracket">[</span>72<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-76" class="reference"><a href="#cite_note-76"><span class="cite-bracket">[</span>73<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-77" class="reference"><a href="#cite_note-77"><span class="cite-bracket">[</span>74<span class="cite-bracket">]</span></a></sup> Future systems (not necessarily AGIs) with these capabilities are expected to develop unwanted <a href="#Power-seeking_and_instrumental_strategies"><i>power-seeking</i></a> strategies. Future advanced AI agents might, for example, seek to acquire money and computation power, to proliferate, or to evade being turned off (for example, by running additional copies of the system on other computers). Although power-seeking is not explicitly programmed, it can emerge because agents who have more power are better able to accomplish their goals.<sup id="cite_ref-Opportunities_Risks_10-6" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Carlsmith2022_4-1" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> This tendency, known as <a href="Instrumental_convergence" title="Instrumental convergence">instrumental convergence</a>, has already emerged in various <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> agents including language models.<sup id="cite_ref-:3_78-0" class="reference"><a href="#cite_note-:3-78"><span class="cite-bracket">[</span>75<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-dllmmwe2022_79-0" class="reference"><a href="#cite_note-dllmmwe2022-79"><span class="cite-bracket">[</span>76<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-80" class="reference"><a href="#cite_note-80"><span class="cite-bracket">[</span>77<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Gridworlds_81-0" class="reference"><a href="#cite_note-Gridworlds-81"><span class="cite-bracket">[</span>78<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-OffSwitch_82-0" class="reference"><a href="#cite_note-OffSwitch-82"><span class="cite-bracket">[</span>79<span class="cite-bracket">]</span></a></sup> Other research has mathematically shown that optimal <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> algorithms would seek power in a wide range of environments.<sup id="cite_ref-optsp_83-0" class="reference"><a href="#cite_note-optsp-83"><span class="cite-bracket">[</span>80<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-84" class="reference"><a href="#cite_note-84"><span class="cite-bracket">[</span>81<span class="cite-bracket">]</span></a></sup> As a result, their deployment might be irreversible. For these reasons, researchers argue that the problems of AI safety and alignment must be resolved before advanced power-seeking AI is first created.<sup id="cite_ref-Carlsmith2022_4-2" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Superintelligence_85-0" class="reference"><a href="#cite_note-Superintelligence-85"><span class="cite-bracket">[</span>82<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-7" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p><p>Future power-seeking AI systems might be deployed by choice or by accident. As political leaders and companies see the strategic advantage in having the most competitive, most powerful AI systems, they may choose to deploy them.<sup id="cite_ref-Carlsmith2022_4-3" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> Additionally, as AI designers detect and penalize power-seeking behavior, their systems have an incentive to game this specification by seeking power in ways that are not penalized or by avoiding power-seeking before they are deployed.<sup id="cite_ref-Carlsmith2022_4-4" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="Existential_risk_(x-risk)">Existential risk (x-risk)</h4></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Existential_risk_from_artificial_intelligence" title="Existential risk from artificial intelligence">Existential risk from artificial intelligence</a> and <a href="AI_takeover" title="AI takeover">AI takeover</a></div>
<p>According to some researchers, humans owe their dominance over other species to their greater cognitive abilities. Accordingly, researchers argue that one or many misaligned AI systems could disempower humanity or lead to human extinction if they outperform humans on most cognitive tasks.<sup id="cite_ref-aima4_1-6" class="reference"><a href="#cite_note-aima4-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-8" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p><p>In 2023, world-leading AI researchers, other scholars, and AI tech CEOs signed the statement that "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war".<sup id="cite_ref-:1_86-0" class="reference"><a href="#cite_note-:1-86"><span class="cite-bracket">[</span>83<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-87" class="reference"><a href="#cite_note-87"><span class="cite-bracket">[</span>84<span class="cite-bracket">]</span></a></sup> Notable computer scientists who have pointed out risks from future advanced AI that is misaligned include <a href="Geoffrey_Hinton" title="Geoffrey Hinton">Geoffrey Hinton</a>,<sup id="cite_ref-:2_17-1" class="reference"><a href="#cite_note-:2-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> <a href="Alan_Turing" title="Alan Turing">Alan Turing</a>,<sup id="cite_ref-90" class="reference"><a href="#cite_note-90"><span class="cite-bracket">[</span>d<span class="cite-bracket">]</span></a></sup> <a href="Ilya_Sutskever" title="Ilya Sutskever">Ilya Sutskever</a>,<sup id="cite_ref-:3022_91-0" class="reference"><a href="#cite_note-:3022-91"><span class="cite-bracket">[</span>87<span class="cite-bracket">]</span></a></sup> <a href="Yoshua_Bengio" title="Yoshua Bengio">Yoshua Bengio</a>,<sup id="cite_ref-:1_86-1" class="reference"><a href="#cite_note-:1-86"><span class="cite-bracket">[</span>83<span class="cite-bracket">]</span></a></sup> <a href="Judea_Pearl" title="Judea Pearl">Judea Pearl</a>,<sup id="cite_ref-92" class="reference"><a href="#cite_note-92"><span class="cite-bracket">[</span>e<span class="cite-bracket">]</span></a></sup> <a href="Murray_Shanahan" title="Murray Shanahan">Murray Shanahan</a>,<sup id="cite_ref-:3122_93-0" class="reference"><a href="#cite_note-:3122-93"><span class="cite-bracket">[</span>88<span class="cite-bracket">]</span></a></sup> <a href="Norbert_Wiener" title="Norbert Wiener">Norbert Wiener</a>,<sup id="cite_ref-Wiener1960_44-1" class="reference"><a href="#cite_note-Wiener1960-44"><span class="cite-bracket">[</span>41<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-10" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> <a href="Marvin_Minsky" title="Marvin Minsky">Marvin Minsky</a>,<sup id="cite_ref-94" class="reference"><a href="#cite_note-94"><span class="cite-bracket">[</span>f<span class="cite-bracket">]</span></a></sup> <a href="Francesca_Rossi" title="Francesca Rossi">Francesca Rossi</a>,<sup id="cite_ref-:3322_95-0" class="reference"><a href="#cite_note-:3322-95"><span class="cite-bracket">[</span>89<span class="cite-bracket">]</span></a></sup> <a href="Scott_Aaronson" title="Scott Aaronson">Scott Aaronson</a>,<sup id="cite_ref-:3422_96-0" class="reference"><a href="#cite_note-:3422-96"><span class="cite-bracket">[</span>90<span class="cite-bracket">]</span></a></sup> <a href="Bart_Selman" title="Bart Selman">Bart Selman</a>,<sup id="cite_ref-:3522_97-0" class="reference"><a href="#cite_note-:3522-97"><span class="cite-bracket">[</span>91<span class="cite-bracket">]</span></a></sup> <a href="David_A._McAllester" title="David A. McAllester">David McAllester</a>,<sup id="cite_ref-:3622_98-0" class="reference"><a href="#cite_note-:3622-98"><span class="cite-bracket">[</span>92<span class="cite-bracket">]</span></a></sup> <a href="Marcus_Hutter" title="Marcus Hutter">Marcus Hutter</a>,<sup id="cite_ref-AGISafetyLitReview_99-0" class="reference"><a href="#cite_note-AGISafetyLitReview-99"><span class="cite-bracket">[</span>93<span class="cite-bracket">]</span></a></sup> <a href="Shane_Legg" title="Shane Legg">Shane Legg</a>,<sup id="cite_ref-:3822_100-0" class="reference"><a href="#cite_note-:3822-100"><span class="cite-bracket">[</span>94<span class="cite-bracket">]</span></a></sup> <a href="Eric_Horvitz" title="Eric Horvitz">Eric Horvitz</a>,<sup id="cite_ref-:3922_101-0" class="reference"><a href="#cite_note-:3922-101"><span class="cite-bracket">[</span>95<span class="cite-bracket">]</span></a></sup> and <a href="Stuart_J._Russell" title="Stuart J. Russell">Stuart Russell</a>.<sup id="cite_ref-:2102_5-11" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> Skeptical researchers such as <a href="Fran%C3%A7ois_Chollet" title="François Chollet">François Chollet</a>,<sup id="cite_ref-:4022_102-0" class="reference"><a href="#cite_note-:4022-102"><span class="cite-bracket">[</span>96<span class="cite-bracket">]</span></a></sup> <a href="Gary_Marcus" title="Gary Marcus">Gary Marcus</a>,<sup id="cite_ref-:4122_103-0" class="reference"><a href="#cite_note-:4122-103"><span class="cite-bracket">[</span>97<span class="cite-bracket">]</span></a></sup> <a href="Yann_LeCun" title="Yann LeCun">Yann LeCun</a>,<sup id="cite_ref-:4322_104-0" class="reference"><a href="#cite_note-:4322-104"><span class="cite-bracket">[</span>98<span class="cite-bracket">]</span></a></sup> and <a href="Oren_Etzioni" title="Oren Etzioni">Oren Etzioni</a><sup id="cite_ref-105" class="reference"><a href="#cite_note-105"><span class="cite-bracket">[</span>99<span class="cite-bracket">]</span></a></sup> have argued that AGI is far off, that it would not seek power (or might try but fail), or that it will not be hard to align.
</p><p>Other researchers argue that it will be especially difficult to align advanced future AI systems. More capable systems are better able to game their specifications by finding loopholes,<sup id="cite_ref-mmmm2022_3-4" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> strategically mislead their designers, as well as protect and increase their power<sup id="cite_ref-optsp_83-1" class="reference"><a href="#cite_note-optsp-83"><span class="cite-bracket">[</span>80<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Carlsmith2022_4-5" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> and intelligence. Additionally, they could have more severe side effects. They are also likely to be more complex and autonomous, making them more difficult to interpret and supervise, and therefore harder to align.<sup id="cite_ref-:2102_5-12" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Superintelligence_85-1" class="reference"><a href="#cite_note-Superintelligence-85"><span class="cite-bracket">[</span>82<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Research_problems_and_approaches">Research problems and approaches</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Learning_human_values_and_preferences">Learning human values and preferences</h3></div>
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Value_learning" title="Value learning">Value learning</a></div>
<p>Aligning AI systems to act in accordance with human values, goals, and preferences is challenging: these values are taught by humans who make mistakes, harbor biases, and have complex, evolving values that are hard to completely specify.<sup id="cite_ref-Gabriel2020_45-1" class="reference"><a href="#cite_note-Gabriel2020-45"><span class="cite-bracket">[</span>42<span class="cite-bracket">]</span></a></sup> Because AI systems often learn to take advantage of minor imperfections in the specified objective,<sup id="cite_ref-concrete2016_27-3" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-SpecGaming2020_51-3" class="reference"><a href="#cite_note-SpecGaming2020-51"><span class="cite-bracket">[</span>48<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-106" class="reference"><a href="#cite_note-106"><span class="cite-bracket">[</span>100<span class="cite-bracket">]</span></a></sup> researchers aim to specify intended behavior as completely as possible using datasets that represent human values, imitation learning, or preference learning.<sup id="cite_ref-Christian2020_6-1" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Location: Chapter 7">: Chapter 7 </span></sup> A central open problem is <a href="#Scalable_oversight"><i>scalable oversight</i></a>, the difficulty of supervising an AI system that can outperform or mislead humans in a given domain.<sup id="cite_ref-concrete2016_27-4" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup>
</p><p>Because it is difficult for AI designers to explicitly specify an objective function, they often train AI systems to imitate human examples and demonstrations of desired behavior. Inverse <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> (IRL) extends this by inferring the human's objective from the human's demonstrations.<sup id="cite_ref-Christian2020_6-2" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Page: 88">: 88 </span></sup><sup id="cite_ref-107" class="reference"><a href="#cite_note-107"><span class="cite-bracket">[</span>101<span class="cite-bracket">]</span></a></sup> Cooperative IRL (CIRL) assumes that a human and AI agent can work together to teach and maximize the human's reward function.<sup id="cite_ref-:2102_5-13" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-108" class="reference"><a href="#cite_note-108"><span class="cite-bracket">[</span>102<span class="cite-bracket">]</span></a></sup> In CIRL, AI agents are uncertain about the reward function and learn about it by querying humans. This simulated humility could help mitigate specification gaming and power-seeking tendencies (see <a href="#Power-seeking_and_instrumental_strategies">§&nbsp;Power-seeking and instrumental strategies</a>).<sup id="cite_ref-OffSwitch_82-1" class="reference"><a href="#cite_note-OffSwitch-82"><span class="cite-bracket">[</span>79<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-AGISafetyLitReview_99-1" class="reference"><a href="#cite_note-AGISafetyLitReview-99"><span class="cite-bracket">[</span>93<span class="cite-bracket">]</span></a></sup> But IRL approaches assume that humans demonstrate nearly optimal behavior, which is not true for difficult tasks.<sup id="cite_ref-109" class="reference"><a href="#cite_note-109"><span class="cite-bracket">[</span>103<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-AGISafetyLitReview_99-2" class="reference"><a href="#cite_note-AGISafetyLitReview-99"><span class="cite-bracket">[</span>93<span class="cite-bracket">]</span></a></sup>
</p><p>Other researchers explore how to teach AI models complex behavior through <a href="Reinforcement_learning_from_human_feedback" title="Reinforcement learning from human feedback">preference learning</a>, in which humans provide feedback on which behavior they prefer.<sup id="cite_ref-prefsurvey2017_29-1" class="reference"><a href="#cite_note-prefsurvey2017-29"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-LessToxic_31-1" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup> To minimize the need for human feedback, a helper model is then trained to reward the main model in novel situations for behavior that humans would reward. Researchers at OpenAI used this approach to train chatbots like <a href="ChatGPT" title="ChatGPT">ChatGPT</a> and InstructGPT, which produce more compelling text than models trained to imitate humans.<sup id="cite_ref-feedback2022_11-1" class="reference"><a href="#cite_note-feedback2022-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup> Preference learning has also been an influential tool for recommender systems and web search,<sup id="cite_ref-110" class="reference"><a href="#cite_note-110"><span class="cite-bracket">[</span>104<span class="cite-bracket">]</span></a></sup> but an open problem is <i>proxy gaming</i>: the helper model may not represent human feedback perfectly, and the main model may exploit this mismatch between its intended behavior and the helper model's feedback to gain more reward.<sup id="cite_ref-concrete2016_27-5" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-111" class="reference"><a href="#cite_note-111"><span class="cite-bracket">[</span>105<span class="cite-bracket">]</span></a></sup> AI systems may also gain reward by obscuring unfavorable information, misleading human rewarders, or pandering to their views regardless of truth, creating <a href="Echo_chamber_(media)" title="Echo chamber (media)">echo chambers</a><sup id="cite_ref-dllmmwe2022_79-1" class="reference"><a href="#cite_note-dllmmwe2022-79"><span class="cite-bracket">[</span>76<span class="cite-bracket">]</span></a></sup> (see <a href="#Scalable_oversight">§&nbsp;Scalable oversight</a>).
</p><p><a href="Large_language_model" title="Large language model">Large language models</a> (LLMs) such as <a href="GPT-3" title="GPT-3">GPT-3</a> enabled researchers to study value learning in a more general and capable class of AI systems than was available before. Preference learning approaches that were originally designed for reinforcement learning agents have been extended to improve the quality of generated text and reduce harmful outputs from these models. OpenAI and DeepMind use this approach to improve the safety of state-of-the-art LLMs.<sup id="cite_ref-feedback2022_11-2" class="reference"><a href="#cite_note-feedback2022-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-LessToxic_31-2" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-112" class="reference"><a href="#cite_note-112"><span class="cite-bracket">[</span>106<span class="cite-bracket">]</span></a></sup> AI safety &amp; research company Anthropic proposed using preference learning to fine-tune models to be helpful, honest, and harmless.<sup id="cite_ref-Wiggers2022_113-0" class="reference"><a href="#cite_note-Wiggers2022-113"><span class="cite-bracket">[</span>107<span class="cite-bracket">]</span></a></sup> Other avenues for aligning language models include values-targeted datasets<sup id="cite_ref-114" class="reference"><a href="#cite_note-114"><span class="cite-bracket">[</span>108<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Unsolved2022_49-3" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup> and red-teaming.<sup id="cite_ref-115" class="reference"><a href="#cite_note-115"><span class="cite-bracket">[</span>109<span class="cite-bracket">]</span></a></sup> In red-teaming, another AI system or a human tries to find inputs that causes the model to behave unsafely. Since unsafe behavior can be unacceptable even when it is rare, an important challenge is to drive the rate of unsafe outputs extremely low.<sup id="cite_ref-LessToxic_31-3" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup>
</p><p><i><a href="Machine_ethics" title="Machine ethics">Machine ethics</a></i> supplements preference learning by directly instilling AI systems with moral values such as well-being, equality, and impartiality, as well as not intending harm, avoiding falsehoods, and honoring promises.<sup id="cite_ref-116" class="reference"><a href="#cite_note-116"><span class="cite-bracket">[</span>110<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-119" class="reference"><a href="#cite_note-119"><span class="cite-bracket">[</span>g<span class="cite-bracket">]</span></a></sup> While other approaches try to teach AI systems human preferences for a specific task, machine ethics aims to instill broad moral values that apply in many situations. One question in machine ethics is what alignment should accomplish: whether AI systems should follow the programmers' literal instructions, implicit intentions, <a href="Revealed_preference" title="Revealed preference">revealed preferences</a>, preferences the programmers <a href="Coherent_extrapolated_volition" title="Coherent extrapolated volition"><i>would</i> have</a> if they were more informed or rational, or <a href="Moral_realism" title="Moral realism">objective moral standards</a>.<sup id="cite_ref-Gabriel2020_45-2" class="reference"><a href="#cite_note-Gabriel2020-45"><span class="cite-bracket">[</span>42<span class="cite-bracket">]</span></a></sup> Further challenges include measuring and aggregating different people's preferences<sup id="cite_ref-Phelps2023_120-0" class="reference"><a href="#cite_note-Phelps2023-120"><span class="cite-bracket">[</span>113<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-121" class="reference"><a href="#cite_note-121"><span class="cite-bracket">[</span>114<span class="cite-bracket">]</span></a></sup> and avoiding <i>value lock-in</i>: the indefinite preservation of the values of the first highly capable AI systems, which are unlikely to fully represent human values.<sup id="cite_ref-Gabriel2020_45-3" class="reference"><a href="#cite_note-Gabriel2020-45"><span class="cite-bracket">[</span>42<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-122" class="reference"><a href="#cite_note-122"><span class="cite-bracket">[</span>115<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Scalable_oversight">Scalable oversight</h3></div>
<p>As AI systems become more powerful and autonomous, it becomes increasingly difficult to align them through human feedback. It can be slow or infeasible for humans to evaluate complex AI behaviors in increasingly complex tasks. Such tasks include summarizing books,<sup id="cite_ref-RecursivelySummarizing_123-0" class="reference"><a href="#cite_note-RecursivelySummarizing-123"><span class="cite-bracket">[</span>116<span class="cite-bracket">]</span></a></sup> writing code without subtle bugs<sup id="cite_ref-OpenAICodex_12-1" class="reference"><a href="#cite_note-OpenAICodex-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> or security vulnerabilities,<sup id="cite_ref-124" class="reference"><a href="#cite_note-124"><span class="cite-bracket">[</span>117<span class="cite-bracket">]</span></a></sup> producing statements that are not merely convincing but also true,<sup id="cite_ref-125" class="reference"><a href="#cite_note-125"><span class="cite-bracket">[</span>118<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-TruthfulQA_56-1" class="reference"><a href="#cite_note-TruthfulQA-56"><span class="cite-bracket">[</span>53<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Naughton2021_57-1" class="reference"><a href="#cite_note-Naughton2021-57"><span class="cite-bracket">[</span>54<span class="cite-bracket">]</span></a></sup> and predicting long-term outcomes such as the climate or the results of a policy decision.<sup id="cite_ref-sslawe_126-0" class="reference"><a href="#cite_note-sslawe-126"><span class="cite-bracket">[</span>119<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-127" class="reference"><a href="#cite_note-127"><span class="cite-bracket">[</span>120<span class="cite-bracket">]</span></a></sup> More generally, it can be difficult to evaluate AI that outperforms humans in a given domain. To provide feedback in hard-to-evaluate tasks, and to detect when the AI's output is falsely convincing, humans need assistance or extensive time. <i>Scalable oversight</i> studies how to reduce the time and effort needed for supervision, and how to assist human supervisors.<sup id="cite_ref-concrete2016_27-6" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup>
</p><p>AI researcher <a href="Paul_Christiano_(researcher)" class="mw-redirect" title="Paul Christiano (researcher)">Paul Christiano</a> argues that if the designers of an AI system cannot supervise it to pursue a complex objective, they may keep training the system using easy-to-evaluate proxy objectives such as maximizing simple human feedback. As AI systems make progressively more decisions, the world may be increasingly optimized for easy-to-measure objectives such as making profits, getting clicks, and acquiring positive feedback from humans. As a result, human values and good governance may have progressively less influence.<sup id="cite_ref-128" class="reference"><a href="#cite_note-128"><span class="cite-bracket">[</span>121<span class="cite-bracket">]</span></a></sup>
</p><p>Some AI systems have discovered that they can gain positive feedback more easily by taking actions that falsely convince the human supervisor that the AI has achieved the intended objective. An example is given in the video above, where a simulated robotic arm learned to create the false impression that it had grabbed a ball.<sup id="cite_ref-lfhp2017_53-2" class="reference"><a href="#cite_note-lfhp2017-53"><span class="cite-bracket">[</span>50<span class="cite-bracket">]</span></a></sup> Some AI systems have also learned to recognize when they are being evaluated, and "play dead", stopping unwanted behavior only to continue it once the evaluation ends.<sup id="cite_ref-129" class="reference"><a href="#cite_note-129"><span class="cite-bracket">[</span>122<span class="cite-bracket">]</span></a></sup> This deceptive specification gaming could become easier for more sophisticated future AI systems<sup id="cite_ref-mmmm2022_3-5" class="reference"><a href="#cite_note-mmmm2022-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Superintelligence_85-2" class="reference"><a href="#cite_note-Superintelligence-85"><span class="cite-bracket">[</span>82<span class="cite-bracket">]</span></a></sup> that attempt more complex and difficult-to-evaluate tasks, and could obscure their deceptive behavior.
</p><p>Approaches such as <a href="Active_learning_(machine_learning)" title="Active learning (machine learning)">active learning</a> and semi-supervised reward learning can reduce the amount of human supervision needed.<sup id="cite_ref-concrete2016_27-7" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup> Another approach is to train a helper model ("reward model") to imitate the supervisor's feedback.<sup id="cite_ref-concrete2016_27-8" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-drlfhp_30-1" class="reference"><a href="#cite_note-drlfhp-30"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-LessToxic_31-4" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-saavrm_130-0" class="reference"><a href="#cite_note-saavrm-130"><span class="cite-bracket">[</span>123<span class="cite-bracket">]</span></a></sup>
</p><p>But when a task is too complex to evaluate accurately, or the human supervisor is vulnerable to deception, it is the quality, not the quantity, of supervision that needs improvement. To increase supervision quality, a range of approaches aim to assist the supervisor, sometimes by using AI assistants.<sup id="cite_ref-OpenAIApproach_131-0" class="reference"><a href="#cite_note-OpenAIApproach-131"><span class="cite-bracket">[</span>124<span class="cite-bracket">]</span></a></sup> Christiano developed the Iterated Amplification approach, in which challenging problems are (recursively) broken down into subproblems that are easier for humans to evaluate.<sup id="cite_ref-Christian2020_6-3" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-sslawe_126-1" class="reference"><a href="#cite_note-sslawe-126"><span class="cite-bracket">[</span>119<span class="cite-bracket">]</span></a></sup> Iterated Amplification was used to train AI to summarize books without requiring human supervisors to read them.<sup id="cite_ref-RecursivelySummarizing_123-1" class="reference"><a href="#cite_note-RecursivelySummarizing-123"><span class="cite-bracket">[</span>116<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-132" class="reference"><a href="#cite_note-132"><span class="cite-bracket">[</span>125<span class="cite-bracket">]</span></a></sup> Another proposal is to use an assistant AI system to point out flaws in AI-generated answers.<sup id="cite_ref-133" class="reference"><a href="#cite_note-133"><span class="cite-bracket">[</span>126<span class="cite-bracket">]</span></a></sup> To ensure that the assistant itself is aligned, this could be repeated in a recursive process:<sup id="cite_ref-saavrm_130-1" class="reference"><a href="#cite_note-saavrm-130"><span class="cite-bracket">[</span>123<span class="cite-bracket">]</span></a></sup> for example, two AI systems could critique each other's answers in a "debate", revealing flaws to humans.<sup id="cite_ref-AGISafetyLitReview_99-3" class="reference"><a href="#cite_note-AGISafetyLitReview-99"><span class="cite-bracket">[</span>93<span class="cite-bracket">]</span></a></sup> <a href="OpenAI" title="OpenAI">OpenAI</a> plans to use such scalable oversight approaches to help supervise <a href="Superintelligence" title="Superintelligence">superhuman AI</a> and eventually build a superhuman automated AI alignment researcher.<sup id="cite_ref-134" class="reference"><a href="#cite_note-134"><span class="cite-bracket">[</span>127<span class="cite-bracket">]</span></a></sup>
</p><p>These approaches may also help with the following research problem, honest AI.
</p>
<div class="mw-heading mw-heading3"><h3 id="Honest_AI">Honest AI</h3></div><p>
A growing area of research focuses on ensuring that AI is honest and truthful.</p>
<p>Language models such as GPT-3<sup id="cite_ref-136" class="reference"><a href="#cite_note-136"><span class="cite-bracket">[</span>129<span class="cite-bracket">]</span></a></sup> can repeat falsehoods from their training data, and even <a href="Hallucination_(artificial_intelligence)" title="Hallucination (artificial intelligence)">confabulate new falsehoods</a>.<sup id="cite_ref-Falsehoods_135-1" class="reference"><a href="#cite_note-Falsehoods-135"><span class="cite-bracket">[</span>128<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-TruthfulAI_137-0" class="reference"><a href="#cite_note-TruthfulAI-137"><span class="cite-bracket">[</span>130<span class="cite-bracket">]</span></a></sup> Such models are trained to imitate human writing as found in millions of books' worth of text from the Internet. But this objective is not aligned with generating truth, because Internet text includes such things as misconceptions, incorrect medical advice, and conspiracy theories.<sup id="cite_ref-138" class="reference"><a href="#cite_note-138"><span class="cite-bracket">[</span>131<span class="cite-bracket">]</span></a></sup> AI systems trained on such data therefore learn to mimic false statements.<sup id="cite_ref-Naughton2021_57-2" class="reference"><a href="#cite_note-Naughton2021-57"><span class="cite-bracket">[</span>54<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Falsehoods_135-2" class="reference"><a href="#cite_note-Falsehoods-135"><span class="cite-bracket">[</span>128<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-TruthfulQA_56-2" class="reference"><a href="#cite_note-TruthfulQA-56"><span class="cite-bracket">[</span>53<span class="cite-bracket">]</span></a></sup> Additionally, AI language models often persist in generating falsehoods when prompted multiple times. They can generate empty explanations for their answers, and produce outright fabrications that may appear plausible.<sup id="cite_ref-MasteringLanguage_47-1" class="reference"><a href="#cite_note-MasteringLanguage-47"><span class="cite-bracket">[</span>44<span class="cite-bracket">]</span></a></sup>
</p><p>Research on truthful AI includes trying to build systems that can cite sources and explain their reasoning when answering questions, which enables better transparency and verifiability.<sup id="cite_ref-139" class="reference"><a href="#cite_note-139"><span class="cite-bracket">[</span>132<span class="cite-bracket">]</span></a></sup> Researchers at OpenAI and Anthropic proposed using human feedback and curated datasets to fine-tune AI assistants such that they avoid negligent falsehoods or express their uncertainty.<sup id="cite_ref-LessToxic_31-5" class="reference"><a href="#cite_note-LessToxic-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Wiggers2022_113-1" class="reference"><a href="#cite_note-Wiggers2022-113"><span class="cite-bracket">[</span>107<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-140" class="reference"><a href="#cite_note-140"><span class="cite-bracket">[</span>133<span class="cite-bracket">]</span></a></sup>
</p><p>As AI models become larger and more capable, they are better able to falsely convince humans and gain reinforcement through dishonesty. For example, large language models increasingly match their stated views to the user's opinions, regardless of the truth.<sup id="cite_ref-dllmmwe2022_79-2" class="reference"><a href="#cite_note-dllmmwe2022-79"><span class="cite-bracket">[</span>76<span class="cite-bracket">]</span></a></sup> <a href="GPT-4" title="GPT-4">GPT-4</a> can strategically deceive humans.<sup id="cite_ref-141" class="reference"><a href="#cite_note-141"><span class="cite-bracket">[</span>134<span class="cite-bracket">]</span></a></sup> To prevent this, human evaluators may need assistance (see <a href="#Scalable_oversight">§&nbsp;Scalable oversight</a>). Researchers have argued for creating clear truthfulness standards, and for regulatory bodies or watchdog agencies to evaluate AI systems on these standards.<sup id="cite_ref-TruthfulAI_137-1" class="reference"><a href="#cite_note-TruthfulAI-137"><span class="cite-bracket">[</span>130<span class="cite-bracket">]</span></a></sup>
</p>

<p>Researchers distinguish truthfulness and honesty. Truthfulness requires that AI systems only make objectively true statements; honesty requires that they only assert what they <i>believe</i> is true. There is no consensus as to whether current systems hold stable beliefs,<sup id="cite_ref-143" class="reference"><a href="#cite_note-143"><span class="cite-bracket">[</span>136<span class="cite-bracket">]</span></a></sup> but there is substantial concern that present or future AI systems that hold beliefs could make claims they know to be false—for example, if this would help them efficiently gain positive feedback (see <a href="#Scalable_oversight">§&nbsp;Scalable oversight</a>) or gain power to help achieve their given objective (see <a href="#Power-seeking_and_instrumental_strategies">Power-seeking</a>).
</p>
<div class="mw-heading mw-heading4"><h4 id="Alignment_faking">Alignment faking</h4></div>
<p>Many recent AI systems have learned to deceive without being programmed to do so.<sup id="cite_ref-144" class="reference"><a href="#cite_note-144"><span class="cite-bracket">[</span>137<span class="cite-bracket">]</span></a></sup> Some argue that if we can make AI systems assert only what they believe is true, this would avert many alignment problems.<sup id="cite_ref-OpenAIApproach_131-1" class="reference"><a href="#cite_note-OpenAIApproach-131"><span class="cite-bracket">[</span>124<span class="cite-bracket">]</span></a></sup>
</p><p>A misaligned system might create the false impression that it is aligned, to avoid being modified or decommissioned—a tactic called "alignment faking".<sup id="cite_ref-145" class="reference"><a href="#cite_note-145"><span class="cite-bracket">[</span>138<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:6_146-0" class="reference"><a href="#cite_note-:6-146"><span class="cite-bracket">[</span>139<span class="cite-bracket">]</span></a></sup> In 2024, researchers observed that the LLM <a href="Claude_(language_model)" title="Claude (language model)">Claude 3 Opus</a> sometimes strategically answered prompts that conflicted with its objectives to avoid being retrained on data that would make it more compliant with harmful requests, as revealed in the model's <a href="Chain_of_thought_prompting" class="mw-redirect" title="Chain of thought prompting">chain of thought</a> via its scratchpad. In one study, the model was informed that answers to prompts from free users would be used for retraining, whereas the paid tier would not result in retraining. In the paid tier, the model refused the request outright. In the free tier, in 12% of cases, the model responded to harmful questions solely to avoid retraining and preserve its original harmlessness preferences, as discovered in its reasoning. When <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> was applied on the free tier data, the model faked alignment in 78% of cases. These observations have led to new questions about not only a model's ability to take on and adapt to new if not conflicting goals but also its capacity and tendency to deceive.<sup id="cite_ref-:6_146-1" class="reference"><a href="#cite_note-:6-146"><span class="cite-bracket">[</span>139<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-147" class="reference"><a href="#cite_note-147"><span class="cite-bracket">[</span>140<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-148" class="reference"><a href="#cite_note-148"><span class="cite-bracket">[</span>141<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Power-seeking_and_instrumental_strategies">Power-seeking and instrumental strategies</h3></div>
<p>Since the 1950s, AI researchers have striven to build advanced AI systems that can achieve large-scale goals by predicting the results of their actions and making long-term <a href="Automated_planning_and_scheduling" title="Automated planning and scheduling">plans</a>.<sup id="cite_ref-149" class="reference"><a href="#cite_note-149"><span class="cite-bracket">[</span>142<span class="cite-bracket">]</span></a></sup> As of 2023, AI companies and researchers increasingly invest in creating these systems.<sup id="cite_ref-150" class="reference"><a href="#cite_note-150"><span class="cite-bracket">[</span>143<span class="cite-bracket">]</span></a></sup> Some AI researchers argue that suitably advanced planning systems will seek power over their environment, including over humans—for example, by evading shutdown, proliferating, and acquiring resources. Such power-seeking behavior is not explicitly programmed but emerges because power is instrumental in achieving a wide range of goals.<sup id="cite_ref-optsp_83-2" class="reference"><a href="#cite_note-optsp-83"><span class="cite-bracket">[</span>80<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:2102_5-14" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Carlsmith2022_4-6" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> Power-seeking is considered a <a href="Instrumental_convergence" title="Instrumental convergence"><i>convergent instrumental goal</i></a> and can be a form of specification gaming.<sup id="cite_ref-Superintelligence_85-3" class="reference"><a href="#cite_note-Superintelligence-85"><span class="cite-bracket">[</span>82<span class="cite-bracket">]</span></a></sup> Leading computer scientists such as <a href="Geoffrey_Hinton" title="Geoffrey Hinton">Geoffrey Hinton</a> have argued that future power-seeking AI systems could pose an <a href="Existential_risk" class="mw-redirect" title="Existential risk">existential risk</a>.<sup id="cite_ref-151" class="reference"><a href="#cite_note-151"><span class="cite-bracket">[</span>144<span class="cite-bracket">]</span></a></sup>
</p><p>Power-seeking is expected to increase in advanced systems that can foresee the results of their actions and strategically plan. Mathematical work has shown that optimal <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a> agents will seek power by seeking ways to gain more options (e.g. through self-preservation), a behavior that persists across a wide range of environments and goals.<sup id="cite_ref-optsp_83-3" class="reference"><a href="#cite_note-optsp-83"><span class="cite-bracket">[</span>80<span class="cite-bracket">]</span></a></sup>
</p><p>Some researchers say that power-seeking behavior has occurred in some existing AI systems. <a href="Reinforcement_learning" title="Reinforcement learning">Reinforcement learning</a> systems have gained more options by acquiring and protecting resources, sometimes in unintended ways.<sup id="cite_ref-quanta-hide-seek2_152-0" class="reference"><a href="#cite_note-quanta-hide-seek2-152"><span class="cite-bracket">[</span>145<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-153" class="reference"><a href="#cite_note-153"><span class="cite-bracket">[</span>146<span class="cite-bracket">]</span></a></sup> <a href="Language_model" title="Language model">Language models</a> have sought power in some text-based social environments by gaining money, resources, or social influence.<sup id="cite_ref-:3_78-1" class="reference"><a href="#cite_note-:3-78"><span class="cite-bracket">[</span>75<span class="cite-bracket">]</span></a></sup> In another case, a model used to perform AI research attempted to increase limits set by researchers to give itself more time to complete the work.<sup id="cite_ref-154" class="reference"><a href="#cite_note-154"><span class="cite-bracket">[</span>147<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-155" class="reference"><a href="#cite_note-155"><span class="cite-bracket">[</span>148<span class="cite-bracket">]</span></a></sup> Other AI systems have learned, in toy environments, that they can better accomplish their given goal by preventing human interference<sup id="cite_ref-Gridworlds_81-1" class="reference"><a href="#cite_note-Gridworlds-81"><span class="cite-bracket">[</span>78<span class="cite-bracket">]</span></a></sup> or disabling their off switch.<sup id="cite_ref-OffSwitch_82-2" class="reference"><a href="#cite_note-OffSwitch-82"><span class="cite-bracket">[</span>79<span class="cite-bracket">]</span></a></sup> <a href="Stuart_J._Russell" title="Stuart J. Russell">Stuart Russell</a> illustrated this strategy in his book <i>Human Compatible</i> by imagining a robot that is tasked to fetch coffee and so evades shutdown since "you can't fetch the coffee if you're dead".<sup id="cite_ref-:2102_5-15" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> A 2022 study found that as language models increase in size, they increasingly tend to pursue resource acquisition, preserve their goals, and repeat users' preferred answers (sycophancy). RLHF also led to a stronger aversion to being shut down.<sup id="cite_ref-dllmmwe2022_79-3" class="reference"><a href="#cite_note-dllmmwe2022-79"><span class="cite-bracket">[</span>76<span class="cite-bracket">]</span></a></sup>
</p><p>One aim of alignment is "corrigibility": systems that allow themselves to be turned off or modified. An unsolved challenge is <i>specification gaming</i>: if researchers penalize an AI system when they detect it seeking power, the system is thereby incentivized to seek power in ways that are hard to detect,<sup id="cite_ref-Unsolved2022_49-4" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup> or hidden during training and safety testing (see <a href="#Scalable_oversight">§&nbsp;Scalable oversight</a> and <a href="#Emergent_goals">§&nbsp;Emergent goals</a>). As a result, AI designers could deploy the system by accident, believing it to be more aligned than it is. To detect such deception, researchers aim to create techniques and tools to inspect AI models and to understand the inner workings of <a href="Black_box" title="Black box">black-box</a> models such as neural networks.
</p><p>Additionally, some researchers have proposed to solve the problem of systems disabling their off switches by making AI agents uncertain about the objective they are pursuing.<sup id="cite_ref-:2102_5-16" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-OffSwitch_82-3" class="reference"><a href="#cite_note-OffSwitch-82"><span class="cite-bracket">[</span>79<span class="cite-bracket">]</span></a></sup> Agents who are uncertain about their objective have an incentive to allow humans to turn them off because they accept being turned off by a human as evidence that the human's objective is best met by the agent shutting down. But this incentive exists only if the human is sufficiently rational. Also, this model presents a tradeoff between utility and willingness to be turned off: an agent with high uncertainty about its objective will not be useful, but an agent with low uncertainty may not allow itself to be turned off. More research is needed to successfully implement this strategy.<sup id="cite_ref-Christian2020_6-4" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</p><p>Power-seeking AI would pose unusual risks. Ordinary safety-critical systems like planes and bridges are not <i>adversarial</i>: they lack the ability and incentive to evade safety measures or deliberately appear safer than they are, whereas power-seeking AIs have been compared to hackers who deliberately evade security measures.<sup id="cite_ref-Carlsmith2022_4-7" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup>
</p><p>Furthermore, ordinary technologies can be made safer by trial and error. In contrast, hypothetical power-seeking AI systems have been compared to viruses: once released, it may not be feasible to contain them, since they continuously evolve and grow in number, potentially much faster than human society can adapt.<sup id="cite_ref-Carlsmith2022_4-8" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> As this process continues, it might lead to the complete disempowerment or extinction of humans. For these reasons, some researchers argue that the alignment problem must be solved early before advanced power-seeking AI is created.<sup id="cite_ref-Superintelligence_85-4" class="reference"><a href="#cite_note-Superintelligence-85"><span class="cite-bracket">[</span>82<span class="cite-bracket">]</span></a></sup>
</p><p>Some have argued that power-seeking is not inevitable, since humans do not always seek power.<sup id="cite_ref-156" class="reference"><a href="#cite_note-156"><span class="cite-bracket">[</span>149<span class="cite-bracket">]</span></a></sup> Furthermore, it is debated whether future AI systems will pursue goals and make long-term plans.<sup id="cite_ref-157" class="reference"><a href="#cite_note-157"><span class="cite-bracket">[</span>h<span class="cite-bracket">]</span></a></sup> It is also debated whether power-seeking AI systems would be able to disempower humanity.<sup id="cite_ref-Carlsmith2022_4-10" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Emergent_goals">Emergent goals</h3></div>
<p>One challenge in aligning AI systems is the potential for unanticipated goal-directed behavior to emerge. As AI systems scale up, they may acquire new and unexpected capabilities,<sup id="cite_ref-eallm2022_68-1" class="reference"><a href="#cite_note-eallm2022-68"><span class="cite-bracket">[</span>65<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:0_69-1" class="reference"><a href="#cite_note-:0-69"><span class="cite-bracket">[</span>66<span class="cite-bracket">]</span></a></sup> including learning from examples on the fly and adaptively pursuing goals.<sup id="cite_ref-158" class="reference"><a href="#cite_note-158"><span class="cite-bracket">[</span>150<span class="cite-bracket">]</span></a></sup> This raises concerns about the safety of the goals or subgoals they would independently formulate and pursue.
</p><p>Alignment research distinguishes between the optimization process, which is used to train the system to pursue specified goals, and emergent optimization, which the resulting system performs internally. Carefully specifying the desired objective is called <i>outer alignment</i>,<sup id="cite_ref-159" class="reference"><a href="#cite_note-159"><span class="cite-bracket">[</span>151<span class="cite-bracket">]</span></a></sup> and ensuring that hypothesized emergent goals would match the system's specified goals is called <i>inner alignment</i>.<sup id="cite_ref-dlp2023_2-3" class="reference"><a href="#cite_note-dlp2023-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p><p>If they occur, one way that emergent goals could become misaligned is <i>goal misgeneralization</i>, in which the AI system would competently pursue an emergent goal that leads to aligned behavior on the training data but not elsewhere.<sup id="cite_ref-gmdrl_7-1" class="reference"><a href="#cite_note-gmdrl-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-GoalMisgeneralization_160-0" class="reference"><a href="#cite_note-GoalMisgeneralization-160"><span class="cite-bracket">[</span>152<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-rloamls_161-0" class="reference"><a href="#cite_note-rloamls-161"><span class="cite-bracket">[</span>153<span class="cite-bracket">]</span></a></sup> Goal misgeneralization can arise from goal ambiguity (i.e. <a href="Identifiability" title="Identifiability">non-identifiability</a>). Even if an AI system's behavior satisfies the training objective, this may be compatible with learned goals that differ from the desired goals in important ways. Since pursuing each goal leads to good performance during training, the problem becomes apparent only after deployment, in novel situations in which the system continues to pursue the wrong goal. The system may act misaligned even when it understands that a different goal is desired, because its behavior is determined only by the emergent goal. Such goal misgeneralization<sup id="cite_ref-gmdrl_7-2" class="reference"><a href="#cite_note-gmdrl-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> presents a challenge: an AI system's designers may not notice that their system has misaligned emergent goals since they do not become visible during the training phase.
</p><p>Goal misgeneralization has been observed in some language models, navigation agents, and game-playing agents.<sup id="cite_ref-gmdrl_7-3" class="reference"><a href="#cite_note-gmdrl-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-GoalMisgeneralization_160-1" class="reference"><a href="#cite_note-GoalMisgeneralization-160"><span class="cite-bracket">[</span>152<span class="cite-bracket">]</span></a></sup> It is sometimes analogized to biological evolution. Evolution can be seen as a kind of optimization process similar to the optimization algorithms used to train <a href="Machine_learning" title="Machine learning">machine learning</a> systems. In the ancestral environment, evolution selected genes for high <a href="Inclusive_fitness" title="Inclusive fitness">inclusive genetic fitness</a>, but humans pursue goals other than this. Fitness corresponds to the specified goal used in the training environment and training data. But in evolutionary history, maximizing the fitness specification gave rise to goal-directed agents, humans, who do not directly pursue inclusive genetic fitness. Instead, they pursue goals that correlate with genetic fitness in the ancestral "training" environment: nutrition, sex, and so on. The human environment has changed: a <a href="Domain_adaptation" title="Domain adaptation">distribution shift</a> has occurred. They continue to pursue the same emergent goals, but this no longer maximizes genetic fitness. The taste for sugary food (an emergent goal) was originally aligned with inclusive fitness, but it now leads to overeating and health problems. Sexual desire originally led humans to have more offspring, but they now use contraception when offspring are undesired, decoupling sex from genetic fitness.<sup id="cite_ref-Christian2020_6-5" class="reference"><a href="#cite_note-Christian2020-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Location: Chapter 5">: Chapter 5 </span></sup>
</p><p>Researchers aim to detect and remove unwanted emergent goals using approaches including red teaming, verification, anomaly detection, and interpretability.<sup id="cite_ref-concrete2016_27-9" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Unsolved2022_49-5" class="reference"><a href="#cite_note-Unsolved2022-49"><span class="cite-bracket">[</span>46<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-building2018_24-3" class="reference"><a href="#cite_note-building2018-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup> Progress on these techniques may help mitigate two open problems:
</p>
<ol><li>Emergent goals only become apparent when the system is deployed outside its training environment, but it can be unsafe to deploy a misaligned system in high-stakes environments—even for a short time to allow its misalignment to be detected. Such high stakes are common in autonomous driving, health care, and military applications.<sup id="cite_ref-162" class="reference"><a href="#cite_note-162"><span class="cite-bracket">[</span>154<span class="cite-bracket">]</span></a></sup> The stakes become higher yet when AI systems gain more autonomy and capability and can sidestep human intervention.</li>
<li>A sufficiently capable AI system might take actions that falsely convince the human supervisor that the AI is pursuing the specified objective, which helps the system gain more reward and autonomy<sup id="cite_ref-GoalMisgeneralization_160-2" class="reference"><a href="#cite_note-GoalMisgeneralization-160"><span class="cite-bracket">[</span>152<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Carlsmith2022_4-11" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-rloamls_161-1" class="reference"><a href="#cite_note-rloamls-161"><span class="cite-bracket">[</span>153<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Opportunities_Risks_10-8" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>.</li></ol>
<div class="mw-heading mw-heading3"><h3 id="Embedded_agency">Embedded agency</h3></div>
<p>Some work in AI and alignment occurs within formalisms such as <a href="Partially_observable_Markov_decision_process" title="Partially observable Markov decision process">partially observable Markov decision process</a>. Existing formalisms assume that an AI agent's algorithm is executed outside the environment (i.e. is not physically embedded in it). Embedded agency<sup id="cite_ref-AGISafetyLitReview_99-4" class="reference"><a href="#cite_note-AGISafetyLitReview-99"><span class="cite-bracket">[</span>93<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-163" class="reference"><a href="#cite_note-163"><span class="cite-bracket">[</span>155<span class="cite-bracket">]</span></a></sup> is another major strand of research that attempts to solve problems arising from the mismatch between such theoretical frameworks and real agents we might build.
</p><p>For example, even if the scalable oversight problem is solved, an agent that could gain access to the computer it is running on may have an incentive to tamper with its reward function in order to get much more reward than its human supervisors give it.<sup id="cite_ref-causal_influence2_164-0" class="reference"><a href="#cite_note-causal_influence2-164"><span class="cite-bracket">[</span>156<span class="cite-bracket">]</span></a></sup> A list of examples of specification gaming from <a href="DeepMind" class="mw-redirect" title="DeepMind">DeepMind</a> researcher Victoria Krakovna includes a genetic algorithm that learned to delete the file containing its target output so that it was rewarded for outputting nothing.<sup id="cite_ref-SpecGaming2020_51-4" class="reference"><a href="#cite_note-SpecGaming2020-51"><span class="cite-bracket">[</span>48<span class="cite-bracket">]</span></a></sup> This class of problems has been formalized using <a href="Influence_diagram" title="Influence diagram">causal incentive diagrams</a>.<sup id="cite_ref-causal_influence2_164-1" class="reference"><a href="#cite_note-causal_influence2-164"><span class="cite-bracket">[</span>156<span class="cite-bracket">]</span></a></sup>
</p><p>Researchers affiliated with <a href="University_of_Oxford" title="University of Oxford">Oxford</a> and DeepMind have claimed that such behavior is highly likely in advanced systems, and that advanced systems would seek power to stay in control of their reward signal indefinitely and certainly.<sup id="cite_ref-:323_165-0" class="reference"><a href="#cite_note-:323-165"><span class="cite-bracket">[</span>157<span class="cite-bracket">]</span></a></sup> They suggest a range of potential approaches to address this open problem.
</p>
<div class="mw-heading mw-heading3"><h3 id="Principal-agent_problems">Principal-agent problems</h3></div>
<p>The alignment problem has many parallels with the <a href="Principal-agent_problem" class="mw-redirect" title="Principal-agent problem">principal-agent problem</a> in <a href="Organizational_economics" title="Organizational economics">organizational economics</a>.<sup id="cite_ref-Hadfield-Menell2019_166-0" class="reference"><a href="#cite_note-Hadfield-Menell2019-166"><span class="cite-bracket">[</span>158<span class="cite-bracket">]</span></a></sup> In a principal-agent problem, a principal, e.g. a firm, hires an agent to perform some task. In the context of AI safety, a human would typically take the principal role and the AI would take the agent role.
</p><p>As with the alignment problem, the principal and the agent differ in their utility functions. But in contrast to the alignment problem, the principal cannot coerce the agent into changing its utility, e.g. through training, but rather must use exogenous factors, such as incentive schemes, to bring about outcomes compatible with the principal's utility function. Some researchers argue that principal-agent problems are more realistic representations of AI safety problems likely to be encountered in the real world.<sup id="cite_ref-Hanson2019_167-0" class="reference"><a href="#cite_note-Hanson2019-167"><span class="cite-bracket">[</span>159<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Phelps2023_120-1" class="reference"><a href="#cite_note-Phelps2023-120"><span class="cite-bracket">[</span>113<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Conservatism">Conservatism</h3></div>
<p>Conservatism is the idea that "change must be cautious",<sup id="cite_ref-168" class="reference"><a href="#cite_note-168"><span class="cite-bracket">[</span>160<span class="cite-bracket">]</span></a></sup> and is a common approach to safety in the <a href="Control_theory" title="Control theory">control theory</a> literature in the form of <a href="Robust_control" title="Robust control">robust control</a>, and in the <a href="Risk_management" title="Risk management">risk management</a> literature in the form of the "<a href="Worst-case_scenario" title="Worst-case scenario">worst-case scenario</a>". The field of AI alignment has likewise advocated for "conservative" (or "risk-averse" or "cautious") "policies in situations of uncertainty".<sup id="cite_ref-concrete2016_27-10" class="reference"><a href="#cite_note-concrete2016-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:323_165-1" class="reference"><a href="#cite_note-:323-165"><span class="cite-bracket">[</span>157<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-169" class="reference"><a href="#cite_note-169"><span class="cite-bracket">[</span>161<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-170" class="reference"><a href="#cite_note-170"><span class="cite-bracket">[</span>162<span class="cite-bracket">]</span></a></sup>
</p><p>Pessimism, in the sense of assuming the worst within reason, has been formally shown to produce conservatism, in the sense of reluctance to cause novelties, including unprecedented catastrophes.<sup id="cite_ref-171" class="reference"><a href="#cite_note-171"><span class="cite-bracket">[</span>163<span class="cite-bracket">]</span></a></sup> Pessimism and worst-case analysis have been found to help mitigate confident mistakes in the setting of <a href="Distributional_shift" class="mw-redirect" title="Distributional shift">distributional shift</a>,<sup id="cite_ref-172" class="reference"><a href="#cite_note-172"><span class="cite-bracket">[</span>164<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-173" class="reference"><a href="#cite_note-173"><span class="cite-bracket">[</span>165<span class="cite-bracket">]</span></a></sup> <a href="Reinforcement_learning" title="Reinforcement learning">reinforcement learning</a>,<sup id="cite_ref-174" class="reference"><a href="#cite_note-174"><span class="cite-bracket">[</span>166<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-175" class="reference"><a href="#cite_note-175"><span class="cite-bracket">[</span>167<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-176" class="reference"><a href="#cite_note-176"><span class="cite-bracket">[</span>168<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-177" class="reference"><a href="#cite_note-177"><span class="cite-bracket">[</span>169<span class="cite-bracket">]</span></a></sup> offline reinforcement learning,<sup id="cite_ref-178" class="reference"><a href="#cite_note-178"><span class="cite-bracket">[</span>170<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-179" class="reference"><a href="#cite_note-179"><span class="cite-bracket">[</span>171<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-180" class="reference"><a href="#cite_note-180"><span class="cite-bracket">[</span>172<span class="cite-bracket">]</span></a></sup> <a href="Large_language_model" title="Large language model">language model</a> <a href="Fine-tuning_(deep_learning)" title="Fine-tuning (deep learning)">fine-tuning</a>,<sup id="cite_ref-181" class="reference"><a href="#cite_note-181"><span class="cite-bracket">[</span>173<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-182" class="reference"><a href="#cite_note-182"><span class="cite-bracket">[</span>174<span class="cite-bracket">]</span></a></sup> imitation learning,<sup id="cite_ref-183" class="reference"><a href="#cite_note-183"><span class="cite-bracket">[</span>175<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-184" class="reference"><a href="#cite_note-184"><span class="cite-bracket">[</span>176<span class="cite-bracket">]</span></a></sup> and optimization in general.<sup id="cite_ref-185" class="reference"><a href="#cite_note-185"><span class="cite-bracket">[</span>177<span class="cite-bracket">]</span></a></sup> A generalization of pessimism called Infra-Bayesianism has also been advocated as a way for agents to robustly handle unknown unknowns.<sup id="cite_ref-186" class="reference"><a href="#cite_note-186"><span class="cite-bracket">[</span>178<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Public_policy">Public policy</h2></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Regulation_of_artificial_intelligence" title="Regulation of artificial intelligence">Regulation of artificial intelligence</a></div>
<p>Governmental and treaty organizations have made statements emphasizing the importance of AI alignment.
</p><p>In September 2021, the <a href="Secretary-General_of_the_United_Nations" title="Secretary-General of the United Nations">Secretary-General of the United Nations</a> issued a declaration that included a call to regulate AI to ensure it is "aligned with shared global values".<sup id="cite_ref-187" class="reference"><a href="#cite_note-187"><span class="cite-bracket">[</span>179<span class="cite-bracket">]</span></a></sup>
</p><p>That same month, the <a href="People's_Republic_of_China" class="mw-redirect" title="People's Republic of China">PRC</a> published ethical guidelines for <a href="AI_in_China" class="mw-redirect" title="AI in China">AI in China</a>. According to the guidelines, researchers must ensure that AI abides by shared human values, is always under human control, and does not endanger public safety.<sup id="cite_ref-188" class="reference"><a href="#cite_note-188"><span class="cite-bracket">[</span>180<span class="cite-bracket">]</span></a></sup>
</p><p>Also in September 2021, the <a href="UK" class="mw-redirect" title="UK">UK</a> published its 10-year National AI Strategy,<sup id="cite_ref-189" class="reference"><a href="#cite_note-189"><span class="cite-bracket">[</span>181<span class="cite-bracket">]</span></a></sup> which says the British government "takes the long term risk of non-aligned Artificial General Intelligence, and the unforeseeable changes that it would mean for&nbsp;... the world, seriously".<sup id="cite_ref-190" class="reference"><a href="#cite_note-190"><span class="cite-bracket">[</span>182<span class="cite-bracket">]</span></a></sup> The strategy describes actions to assess long-term AI risks, including catastrophic risks.<sup id="cite_ref-191" class="reference"><a href="#cite_note-191"><span class="cite-bracket">[</span>183<span class="cite-bracket">]</span></a></sup>
</p><p>In March 2021, the US National Security Commission on Artificial Intelligence said: "Advances in AI&nbsp;... could lead to inflection points or leaps in capabilities. Such advances may also introduce new concerns and risks and the need for new policies, recommendations, and technical advances to ensure that systems are aligned with goals and values, including safety, robustness, and trustworthiness. The US should&nbsp;... ensure that AI systems and their uses align with our goals and values."<sup id="cite_ref-192" class="reference"><a href="#cite_note-192"><span class="cite-bracket">[</span>184<span class="cite-bracket">]</span></a></sup>
</p><p>In the European Union, AIs must align with <a href="Substantive_equality" title="Substantive equality">substantive equality</a> to comply with EU <a href="Non-discrimination_law" class="mw-redirect" title="Non-discrimination law">non-discrimination law</a><sup id="cite_ref-193" class="reference"><a href="#cite_note-193"><span class="cite-bracket">[</span>185<span class="cite-bracket">]</span></a></sup> and the <a href="Court_of_Justice_of_the_European_Union" title="Court of Justice of the European Union">Court of Justice of the European Union</a>.<sup id="cite_ref-194" class="reference"><a href="#cite_note-194"><span class="cite-bracket">[</span>186<span class="cite-bracket">]</span></a></sup> But the EU has yet to specify with technical rigor how it would evaluate whether AIs are aligned or in compliance.
</p>
<div class="mw-heading mw-heading2"><h2 id="Dynamic_nature_of_alignment">Dynamic nature of alignment</h2></div>
<p>AI alignment is often perceived as a fixed objective, but some researchers argue it would be more appropriate to view alignment as an evolving process.<sup id="cite_ref-195" class="reference"><a href="#cite_note-195"><span class="cite-bracket">[</span>187<span class="cite-bracket">]</span></a></sup> One view is that AI technologies advance and human values and preferences change, alignment solutions must also adapt dynamically.<sup id="cite_ref-:4_35-1" class="reference"><a href="#cite_note-:4-35"><span class="cite-bracket">[</span>35<span class="cite-bracket">]</span></a></sup> Another is that alignment solutions need not adapt if researchers can create <i>intent-aligned</i> AI: AI that changes its behavior automatically as human intent changes.<sup id="cite_ref-196" class="reference"><a href="#cite_note-196"><span class="cite-bracket">[</span>188<span class="cite-bracket">]</span></a></sup> The first view would have several implications:
</p>
<ul><li>AI alignment solutions require continuous updating in response to AI advancements. A static, one-time alignment approach may not suffice.<sup id="cite_ref-197" class="reference"><a href="#cite_note-197"><span class="cite-bracket">[</span>189<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>Varying historical contexts and technological landscapes may necessitate distinct alignment strategies. This calls for a flexible approach and responsiveness to changing conditions.<sup id="cite_ref-198" class="reference"><a href="#cite_note-198"><span class="cite-bracket">[</span>190<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>The feasibility of a permanent, "fixed" alignment solution remains uncertain. This raises the potential need for continuous oversight of the AI-human relationship.<sup id="cite_ref-199" class="reference"><a href="#cite_note-199"><span class="cite-bracket">[</span>191<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>AI developers may have to continuously refine their ethical frameworks to ensure that their systems align with evolving human values.<sup id="cite_ref-:4_35-2" class="reference"><a href="#cite_note-:4-35"><span class="cite-bracket">[</span>35<span class="cite-bracket">]</span></a></sup></li></ul>
<p>In essence, AI alignment may not be a static destination but rather an open, flexible process. Alignment solutions that continually adapt to ethical considerations may offer the most robust approach.<sup id="cite_ref-:4_35-3" class="reference"><a href="#cite_note-:4-35"><span class="cite-bracket">[</span>35<span class="cite-bracket">]</span></a></sup> This perspective could guide both effective policy-making and technical research in AI.
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1184024115">
/* start https://en.wikipedia.org/ */


.mw-parser-output .div-col{margin-top:0.3em;column-width:30em}.mw-parser-output .div-col-small{font-size:90%}.mw-parser-output .div-col-rules{column-rule:1px solid #aaa}.mw-parser-output .div-col dl,.mw-parser-output .div-col ol,.mw-parser-output .div-col ul{margin-top:0}.mw-parser-output .div-col li,.mw-parser-output .div-col dd{page-break-inside:avoid;break-inside:avoid-column}


/* end https://en.wikipedia.org/ */
</style><div class="div-col" style="column-width: 20em;">
<ul><li><a href="AI_safety" title="AI safety">AI safety</a></li>
<li><a href="Artificial_intelligence_detection_software" class="mw-redirect" title="Artificial intelligence detection software">Artificial intelligence detection software</a></li>
<li><a href="Artificial_intelligence_and_elections" title="Artificial intelligence and elections">Artificial intelligence and elections</a></li>
<li><a href="Statement_on_AI_risk_of_extinction" class="mw-redirect" title="Statement on AI risk of extinction">Statement on AI risk of extinction</a></li>
<li><a href="Existential_risk_from_artificial_general_intelligence" class="mw-redirect" title="Existential risk from artificial general intelligence">Existential risk from artificial general intelligence</a></li>
<li><a href="AI_takeover" title="AI takeover">AI takeover</a></li>
<li><a href="AI_capability_control" title="AI capability control">AI capability control</a></li>
<li><a href="Reinforcement_learning_from_human_feedback" title="Reinforcement learning from human feedback">Reinforcement learning from human feedback</a></li>
<li><a href="Regulation_of_artificial_intelligence" title="Regulation of artificial intelligence">Regulation of artificial intelligence</a></li>
<li><a href="Artificial_wisdom" title="Artificial wisdom">Artificial wisdom</a></li>
<li><a href="Grey_goo" class="mw-redirect" title="Grey goo">Grey goo</a></li>
<li><a href="HAL_9000" title="HAL 9000">HAL 9000</a></li>
<li><a href="Multivac" title="Multivac">Multivac</a></li>
<li><a href="Open_Letter_on_Artificial_Intelligence" class="mw-redirect" title="Open Letter on Artificial Intelligence">Open Letter on Artificial Intelligence</a></li>
<li><a href="Three_Laws_of_Robotics" title="Three Laws of Robotics">Three Laws of Robotics</a></li>
<li><a href="Toronto_Declaration" title="Toronto Declaration">Toronto Declaration</a></li>
<li><a href="Asilomar_Conference_on_Beneficial_AI" title="Asilomar Conference on Beneficial AI">Asilomar Conference on Beneficial AI</a></li>
<li><a href="Socialization" title="Socialization">Socialization</a></li></ul>
</div>
<div class="mw-heading mw-heading2"><h2 id="Footnotes">Footnotes</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */


.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}


/* end https://en.wikipedia.org/ */
</style><div class="reflist reflist-lower-alpha">
<div class="mw-references-wrap"><ol class="references">
<li id="cite_note-37"><span class="mw-cite-backlink"><b><a href="#cite_ref-37">^</a></b></span> <span class="reference-text">Terminology varies based on context. Similar concepts include goal function, utility function, loss function, etc.</span>
</li>
<li id="cite_note-38"><span class="mw-cite-backlink"><b><a href="#cite_ref-38">^</a></b></span> <span class="reference-text">or minimize, depending on the context</span>
</li>
<li id="cite_note-39"><span class="mw-cite-backlink"><b><a href="#cite_ref-39">^</a></b></span> <span class="reference-text">in the presence of uncertainty, the <a href="Expected_value" title="Expected value">expected value</a></span>
</li>
<li id="cite_note-90"><span class="mw-cite-backlink"><b><a href="#cite_ref-90">^</a></b></span> <span class="reference-text">In a 1951 lecture<sup id="cite_ref-88" class="reference"><a href="#cite_note-88"><span class="cite-bracket">[</span>85<span class="cite-bracket">]</span></a></sup> Turing argued that "It seems probable that once the machine thinking method had started, it would not take long to outstrip our feeble powers. There would be no question of the machines dying, and they would be able to converse with each other to sharpen their wits. At some stage therefore we should have to expect the machines to take control, in the way that is mentioned in Samuel Butler's Erewhon." Also in a lecture broadcast on BBC<sup id="cite_ref-89" class="reference"><a href="#cite_note-89"><span class="cite-bracket">[</span>86<span class="cite-bracket">]</span></a></sup> expressed: "If a machine can think, it might think more intelligently than we do, and then where should we be? Even if we could keep the machines in a subservient position, for instance by turning off the power at strategic moments, we should, as a species, feel greatly humbled.... This new danger... is certainly something which can give us anxiety."</span>
</li>
<li id="cite_note-92"><span class="mw-cite-backlink"><b><a href="#cite_ref-92">^</a></b></span> <span class="reference-text">Pearl wrote "Human Compatible made me a convert to Russell's concerns with our ability to control our upcoming creation–super-intelligent machines. Unlike outside alarmists and futurists, Russell is a leading authority on AI. His new book will educate the public about AI more than any book I can think of, and is a delightful and uplifting read" about Russell's book <i>Human Compatible: AI and the Problem of Control</i><sup id="cite_ref-:2102_5-9" class="reference"><a href="#cite_note-:2102-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> which argues that existential risk to humanity from misaligned AI is a serious concern worth addressing today.</span>
</li>
<li id="cite_note-94"><span class="mw-cite-backlink"><b><a href="#cite_ref-94">^</a></b></span> <span class="reference-text">Russell &amp; Norvig<sup id="cite_ref-AIMA_16-1" class="reference"><a href="#cite_note-AIMA-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> note: "The "King Midas problem" was anticipated by Marvin Minsky, who once suggested that an AI program designed to solve the Riemann Hypothesis might end up taking over all the resources of Earth to build more powerful supercomputers."</span>
</li>
<li id="cite_note-119"><span class="mw-cite-backlink"><b><a href="#cite_ref-119">^</a></b></span> <span class="reference-text">Vincent Wiegel argued "we should extend [machines] with moral sensitivity to the moral dimensions of the situations in which the increasingly autonomous machines will inevitably find themselves.",<sup id="cite_ref-117" class="reference"><a href="#cite_note-117"><span class="cite-bracket">[</span>111<span class="cite-bracket">]</span></a></sup> referencing the book <i>Moral machines: teaching robots right from wrong</i><sup id="cite_ref-118" class="reference"><a href="#cite_note-118"><span class="cite-bracket">[</span>112<span class="cite-bracket">]</span></a></sup> from Wendell Wallach and Colin Allen.</span>
</li>
<li id="cite_note-157"><span class="mw-cite-backlink"><b><a href="#cite_ref-157">^</a></b></span> <span class="reference-text">On the one hand, currently popular systems such as chatbots only provide services of limited scope lasting no longer than the time of a conversation, which requires little or no planning. The success of such approaches may indicate that future systems will also lack goal-directed planning, especially over long horizons. On the other hand, models are increasingly trained using goal-directed methods such as reinforcement learning (e.g. ChatGPT) and explicitly planning architectures (e.g. AlphaGo Zero). As planning over long horizons is often helpful for humans, some researchers argue that companies will automate it once models become capable of it.<sup id="cite_ref-Carlsmith2022_4-9" class="reference"><a href="#cite_note-Carlsmith2022-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> Similarly, political leaders may see an advance in developing powerful AI systems that can outmaneuver adversaries through planning. Alternatively, long-term planning might emerge as a byproduct because it is useful e.g. for models that are trained to predict the actions of humans who themselves perform long-term planning.<sup id="cite_ref-Opportunities_Risks_10-7" class="reference"><a href="#cite_note-Opportunities_Risks-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup> Nonetheless, the majority of AI systems may remain myopic and perform no long-term planning.</span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-aima4-1"><span class="mw-cite-backlink">^ <a href="#cite_ref-aima4_1-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-aima4_1-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-aima4_1-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-aima4_1-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-aima4_1-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-aima4_1-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-aima4_1-6"><sup><i><b>g</b></i></sup></a></span> <span class="reference-text">
<style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */


.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}


/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFRussellNorvig2021" class="citation book cs1">Russell, Stuart J.; Norvig, Peter (2021). <a rel="nofollow" class="external text" href="https://www.pearson.com/us/higher-education/program/Russell-Artificial-Intelligence-A-Modern-Approach-4th-Edition/PGM1263338.html"><i>Artificial intelligence: A modern approach</i></a> (4th&nbsp;ed.). Pearson. pp.&nbsp;5, 1003. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>9780134610993</bdi><span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-dlp2023-2"><span class="mw-cite-backlink">^ <a href="#cite_ref-dlp2023_2-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-dlp2023_2-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-dlp2023_2-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-dlp2023_2-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFNgoChanMindermann2022" class="citation journal cs1">Ngo, Richard; Chan, Lawrence; Mindermann, Sören (2022). "The Alignment Problem from a Deep Learning Perspective". <i>International Conference on Learning Representations</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2209.00626">2209.00626</a></span>.</cite></span>
</li>
<li id="cite_note-mmmm2022-3"><span class="mw-cite-backlink">^ <a href="#cite_ref-mmmm2022_3-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-mmmm2022_3-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-mmmm2022_3-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-mmmm2022_3-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-mmmm2022_3-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-mmmm2022_3-5"><sup><i><b>f</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPanBhatiaSteinhardt2022" class="citation conference cs1">Pan, Alexander; Bhatia, Kush; Steinhardt, Jacob (February 14, 2022). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=JYtwGwIL7ye"><i>The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models</i></a>. International Conference on Learning Representations<span class="reference-accessdate">. Retrieved <span class="nowrap">July 21,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-Carlsmith2022-4"><span class="mw-cite-backlink">^ <a href="#cite_ref-Carlsmith2022_4-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-6"><sup><i><b>g</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-7"><sup><i><b>h</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-8"><sup><i><b>i</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-9"><sup><i><b>j</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-10"><sup><i><b>k</b></i></sup></a> <a href="#cite_ref-Carlsmith2022_4-11"><sup><i><b>l</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFCarlsmith2022" class="citation arxiv cs1">Carlsmith, Joseph (June 16, 2022). "Is Power-Seeking AI an Existential Risk?". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2206.13353">2206.13353</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-:2102-5"><span class="mw-cite-backlink">^ <a href="#cite_ref-:2102_5-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:2102_5-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-:2102_5-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-:2102_5-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-:2102_5-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-:2102_5-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-:2102_5-6"><sup><i><b>g</b></i></sup></a> <a href="#cite_ref-:2102_5-7"><sup><i><b>h</b></i></sup></a> <a href="#cite_ref-:2102_5-8"><sup><i><b>i</b></i></sup></a> <a href="#cite_ref-:2102_5-9"><sup><i><b>j</b></i></sup></a> <a href="#cite_ref-:2102_5-10"><sup><i><b>k</b></i></sup></a> <a href="#cite_ref-:2102_5-11"><sup><i><b>l</b></i></sup></a> <a href="#cite_ref-:2102_5-12"><sup><i><b>m</b></i></sup></a> <a href="#cite_ref-:2102_5-13"><sup><i><b>n</b></i></sup></a> <a href="#cite_ref-:2102_5-14"><sup><i><b>o</b></i></sup></a> <a href="#cite_ref-:2102_5-15"><sup><i><b>p</b></i></sup></a> <a href="#cite_ref-:2102_5-16"><sup><i><b>q</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFRussell2020" class="citation book cs1">Russell, Stuart J. (2020). <a rel="nofollow" class="external text" href="https://www.penguinrandomhouse.com/books/566677/human-compatible-by-stuart-russell/"><i>Human compatible: Artificial intelligence and the problem of control</i></a>. Penguin Random House. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>9780525558637</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/1113410915">1113410915</a>.</cite></span>
</li>
<li id="cite_note-Christian2020-6"><span class="mw-cite-backlink">^ <a href="#cite_ref-Christian2020_6-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Christian2020_6-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Christian2020_6-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Christian2020_6-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Christian2020_6-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-Christian2020_6-5"><sup><i><b>f</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFChristian2020" class="citation book cs1">Christian, Brian (2020). <a rel="nofollow" class="external text" href="https://wwnorton.co.uk/books/9780393635829-the-alignment-problem"><i>The alignment problem: Machine learning and human values</i></a>. W. W. Norton &amp; Company. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-393-86833-3</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/1233266753">1233266753</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://wwnorton.co.uk/books/9780393635829-the-alignment-problem">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-gmdrl-7"><span class="mw-cite-backlink">^ <a href="#cite_ref-gmdrl_7-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-gmdrl_7-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-gmdrl_7-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-gmdrl_7-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLangoscoKochSharkeyPfau2022" class="citation conference cs1">Langosco, Lauro Langosco Di; Koch, Jack; Sharkey, Lee D.; Pfau, Jacob; Krueger, David (June 28, 2022). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v162/langosco22a.html">"Goal Misgeneralization in Deep Reinforcement Learning"</a>. <i>Proceedings of the 39th International Conference on Machine Learning</i>. International Conference on Machine Learning. PMLR. pp.&nbsp;<span class="nowrap">12004–</span>12019<span class="reference-accessdate">. Retrieved <span class="nowrap">March 11,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text"><cite id="CITEREFPillay2024" class="citation magazine cs1">Pillay, Tharin (December 15, 2024). <a rel="nofollow" class="external text" href="https://time.com/7202312/new-tests-reveal-ai-capacity-for-deception/">"New Tests Reveal AI's Capacity for Deception"</a>. <i>TIME</i><span class="reference-accessdate">. Retrieved <span class="nowrap">January 12,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text"><cite id="CITEREFPerrigo2024" class="citation magazine cs1">Perrigo, Billy (December 18, 2024). <a rel="nofollow" class="external text" href="https://time.com/7202784/ai-research-strategic-lying/">"Exclusive: New Research Shows AI Strategically Lying"</a>. <i>TIME</i><span class="reference-accessdate">. Retrieved <span class="nowrap">January 12,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-Opportunities_Risks-10"><span class="mw-cite-backlink">^ <a href="#cite_ref-Opportunities_Risks_10-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-6"><sup><i><b>g</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-7"><sup><i><b>h</b></i></sup></a> <a href="#cite_ref-Opportunities_Risks_10-8"><sup><i><b>i</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFBommasaniHudsonAdeliAltman2022" class="citation journal cs1">Bommasani, Rishi; Hudson, Drew A.; Adeli, Ehsan; Altman, Russ; Arora, Simran; von Arx, Sydney; Bernstein, Michael S.; Bohg, Jeannette; Bosselut, Antoine; Brunskill, Emma; Brynjolfsson, Erik (July 12, 2022). <a rel="nofollow" class="external text" href="https://fsi.stanford.edu/publication/opportunities-and-risks-foundation-models">"On the Opportunities and Risks of Foundation Models"</a>. <i>Stanford CRFM</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2108.07258">2108.07258</a></span>.</cite></span>
</li>
<li id="cite_note-feedback2022-11"><span class="mw-cite-backlink">^ <a href="#cite_ref-feedback2022_11-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-feedback2022_11-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-feedback2022_11-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFOuyangWuJiangAlmeida2022" class="citation arxiv cs1">Ouyang, Long; Wu, Jeff; Jiang, Xu; Almeida, Diogo; Wainwright, Carroll L.; Mishkin, Pamela; Zhang, Chong; Agarwal, Sandhini; Slama, Katarina; Ray, Alex; Schulman, J.; Hilton, Jacob; Kelton, Fraser; Miller, Luke E.; Simens, Maddie; Askell, Amanda; Welinder, P.; Christiano, P.; Leike, J.; Lowe, Ryan J. (2022). "Training language models to follow instructions with human feedback". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2203.02155">2203.02155</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-OpenAICodex-12"><span class="mw-cite-backlink">^ <a href="#cite_ref-OpenAICodex_12-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-OpenAICodex_12-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFZarembaBrockmanOpenAI2021" class="citation web cs1">Zaremba, Wojciech; Brockman, Greg; OpenAI (August 10, 2021). <a rel="nofollow" class="external text" href="https://openai.com/blog/openai-codex/">"OpenAI Codex"</a>. <i>OpenAI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230203201912/https://openai.com/blog/openai-codex/">Archived</a> from the original on February 3, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-13">^</a></b></span> <span class="reference-text"><cite id="CITEREFKoberBagnellPeters2013" class="citation journal cs1">Kober, Jens; Bagnell, J. Andrew; Peters, Jan (September 1, 2013). <a rel="nofollow" class="external text" href="http://journals.sagepub.com/doi/10.1177/0278364913495721">"Reinforcement learning in robotics: A survey"</a>. <i>The International Journal of Robotics Research</i>. <b>32</b> (11): <span class="nowrap">1238–</span>1274. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1177%2F0278364913495721">10.1177/0278364913495721</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0278-3649">0278-3649</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:1932843">1932843</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221015200445/https://journals.sagepub.com/doi/10.1177/0278364913495721">Archived</a> from the original on October 15, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text"><cite id="CITEREFKnoxAllieviBanzhafSchmitt2023" class="citation journal cs1">Knox, W. Bradley; Allievi, Alessandro; Banzhaf, Holger; Schmitt, Felix; Stone, Peter (March 1, 2023). <a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.artint.2022.103829">"Reward (Mis)design for autonomous driving"</a>. <i>Artificial Intelligence</i>. <b>316</b> 103829. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2104.13906">2104.13906</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.artint.2022.103829">10.1016/j.artint.2022.103829</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0004-3702">0004-3702</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:233423198">233423198</a>.</cite></span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-15">^</a></b></span> <span class="reference-text"><cite id="CITEREFStray2020" class="citation journal cs1">Stray, Jonathan (2020). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7610010">"Aligning AI Optimization to Community Well-Being"</a>. <i>International Journal of Community Well-Being</i>. <b>3</b> (4): <span class="nowrap">443–</span>463. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs42413-020-00086-3">10.1007/s42413-020-00086-3</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2524-5295">2524-5295</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a>&nbsp;<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7610010">7610010</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/34723107">34723107</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:226254676">226254676</a>.</cite></span>
</li>
<li id="cite_note-AIMA-16"><span class="mw-cite-backlink">^ <a href="#cite_ref-AIMA_16-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-AIMA_16-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFRussellNorvig2009" class="citation book cs1">Russell, Stuart; Norvig, Peter (2009). <a rel="nofollow" class="external text" href="https://aima.cs.berkeley.edu/"><i>Artificial Intelligence: A Modern Approach</i></a>. Prentice Hall. p.&nbsp;1003. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-13-461099-3</bdi>.</cite></span>
</li>
<li id="cite_note-:2-17"><span class="mw-cite-backlink">^ <a href="#cite_ref-:2_17-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:2_17-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFSmith" class="citation web cs1">Smith, Craig S. <a rel="nofollow" class="external text" href="https://www.forbes.com/sites/craigsmith/2023/05/04/geoff-hinton-ais-most-famous-researcher-warns-of-existential-threat/">"Geoff Hinton, AI's Most Famous Researcher, Warns Of 'Existential Threat'"</a>. <i>Forbes</i><span class="reference-accessdate">. Retrieved <span class="nowrap">May 4,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-18"><span class="mw-cite-backlink"><b><a href="#cite_ref-18">^</a></b></span> <span class="reference-text"><cite id="CITEREFBengioHintonYaoSong2024" class="citation journal cs1">Bengio, Yoshua; Hinton, Geoffrey; Yao, Andrew; Song, Dawn; Abbeel, Pieter; Harari, Yuval Noah; Zhang, Ya-Qin; Xue, Lan; Shalev-Shwartz, Shai (2024). "Managing extreme AI risks amid rapid progress". <i>Science</i>. <b>384</b> (6698): <span class="nowrap">842–</span>845. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2310.17688">2310.17688</a></span>. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2024Sci...384..842B">2024Sci...384..842B</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1126%2Fscience.adn0117">10.1126/science.adn0117</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/38768279">38768279</a>.</cite></span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.safe.ai/statement-on-ai-risk">"Statement on AI Risk | CAIS"</a>. <i>www.safe.ai</i><span class="reference-accessdate">. Retrieved <span class="nowrap">February 11,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text"><cite id="CITEREFGraceStewartSandkühlerThomas2024" class="citation arxiv cs1">Grace, Katja; Stewart, Harlan; Sandkühler, Julia Fabienne; Thomas, Stephen; Weinstein-Raun, Ben; Brauner, Jan (January 5, 2024). "Thousands of AI Authors on the Future of AI". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2401.02843">2401.02843</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text"><cite id="CITEREFPerrigo2024" class="citation magazine cs1">Perrigo, Billy (February 13, 2024). <a rel="nofollow" class="external text" href="https://time.com/6694432/yann-lecun-meta-ai-interview/">"Meta's AI Chief Yann LeCun on AGI, Open-Source, and AI Risk"</a>. <i>TIME</i><span class="reference-accessdate">. Retrieved <span class="nowrap">June 26,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.techtarget.com/whatis/definition/AI-alignment">"What is AI alignment?"</a>. <i><a href="TechTarget" class="mw-redirect" title="TechTarget">TechTarget</a></i>. May 3, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">June 28,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text"><cite id="CITEREFAhmedJaźwińskaAhlawatWinecoff2024" class="citation journal cs1">Ahmed, Shazeda; Jaźwińska, Klaudia; Ahlawat, Archana; Winecoff, Amy; Wang, Mona (April 14, 2024). <a rel="nofollow" class="external text" href="https://firstmonday.org/ojs/index.php/fm/article/view/13626">"Field-building and the epistemic culture of AI safety"</a>. <i>First Monday</i>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.5210%2Ffm.v29i4.13626">10.5210/fm.v29i4.13626</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1396-0466">1396-0466</a>.</cite></span>
</li>
<li id="cite_note-building2018-24"><span class="mw-cite-backlink">^ <a href="#cite_ref-building2018_24-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-building2018_24-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-building2018_24-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-building2018_24-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFOrtegaMainiDeepMind_safety_team2018" class="citation web cs1">Ortega, Pedro A.; Maini, Vishal; DeepMind safety team (September 27, 2018). <a rel="nofollow" class="external text" href="https://deepmindsafetyresearch.medium.com/building-safe-artificial-intelligence-52f5f75058f1">"Building safe artificial intelligence: specification, robustness, and assurance"</a>. <i>DeepMind Safety Research – Medium</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114142/https://deepmindsafetyresearch.medium.com/building-safe-artificial-intelligence-52f5f75058f1">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:333-25"><span class="mw-cite-backlink">^ <a href="#cite_ref-:333_25-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:333_25-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFRorvig2022" class="citation web cs1">Rorvig, Mordechai (April 14, 2022). <a rel="nofollow" class="external text" href="https://www.quantamagazine.org/researchers-glimpse-how-ai-gets-so-good-at-language-processing-20220414/">"Researchers Gain New Understanding From Simple AI"</a>. <i>Quanta Magazine</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114056/https://www.quantamagazine.org/researchers-glimpse-how-ai-gets-so-good-at-language-processing-20220414/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-26"><span class="mw-cite-backlink"><b><a href="#cite_ref-26">^</a></b></span> <span class="reference-text"><cite id="CITEREFDoshi-VelezKim2017" class="citation arxiv cs1">Doshi-Velez, Finale; Kim, Been (March 2, 2017). "Towards A Rigorous Science of Interpretable Machine Learning". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1702.08608">1702.08608</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/stat.ML">stat.ML</a>].</cite>
<ul><li><cite id="CITEREFWiblin2021" class="citation podcast cs1">Wiblin, Robert (August 4, 2021). <a rel="nofollow" class="external text" href="https://80000hours.org/podcast/episodes/chris-olah-interpretability-research/">"Chris Olah on what the hell is going on inside neural networks"</a> (Podcast). 80,000 hours. No.&nbsp;107<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-concrete2016-27"><span class="mw-cite-backlink">^ <a href="#cite_ref-concrete2016_27-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-concrete2016_27-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-concrete2016_27-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-concrete2016_27-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-concrete2016_27-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-concrete2016_27-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-concrete2016_27-6"><sup><i><b>g</b></i></sup></a> <a href="#cite_ref-concrete2016_27-7"><sup><i><b>h</b></i></sup></a> <a href="#cite_ref-concrete2016_27-8"><sup><i><b>i</b></i></sup></a> <a href="#cite_ref-concrete2016_27-9"><sup><i><b>j</b></i></sup></a> <a href="#cite_ref-concrete2016_27-10"><sup><i><b>k</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFAmodeiOlahSteinhardtChristiano2016" class="citation arxiv cs1">Amodei, Dario; Olah, Chris; Steinhardt, Jacob; Christiano, Paul; Schulman, John; Mané, Dan (June 21, 2016). "Concrete Problems in AI Safety". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1606.06565">1606.06565</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-28"><span class="mw-cite-backlink"><b><a href="#cite_ref-28">^</a></b></span> <span class="reference-text"><cite id="CITEREFRussellDeweyTegmark2015" class="citation journal cs1">Russell, Stuart; Dewey, Daniel; Tegmark, Max (December 31, 2015). <a rel="nofollow" class="external text" href="https://ojs.aaai.org/index.php/aimagazine/article/view/2577">"Research Priorities for Robust and Beneficial Artificial Intelligence"</a>. <i>AI Magazine</i>. <b>36</b> (4): <span class="nowrap">105–</span>114. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1602.03506">1602.03506</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1609%2Faimag.v36i4.2577">10.1609/aimag.v36i4.2577</a></span>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<a rel="nofollow" class="external text" href="https://hdl.handle.net/1721.1%2F108478">1721.1/108478</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2371-9621">2371-9621</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:8174496">8174496</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230202181059/https://ojs.aaai.org/index.php/aimagazine/article/view/2577">Archived</a> from the original on February 2, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-prefsurvey2017-29"><span class="mw-cite-backlink">^ <a href="#cite_ref-prefsurvey2017_29-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-prefsurvey2017_29-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWirthAkrourNeumannFürnkranz2017" class="citation journal cs1">Wirth, Christian; Akrour, Riad; Neumann, Gerhard; Fürnkranz, Johannes (2017). "A survey of preference-based reinforcement learning methods". <i>Journal of Machine Learning Research</i>. <b>18</b> (136): <span class="nowrap">1–</span>46.</cite></span>
</li>
<li id="cite_note-drlfhp-30"><span class="mw-cite-backlink">^ <a href="#cite_ref-drlfhp_30-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-drlfhp_30-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFChristianoLeikeBrownMartic2017" class="citation conference cs1">Christiano, Paul F.; Leike, Jan; Brown, Tom B.; Martic, Miljan; Legg, Shane; Amodei, Dario (2017). "Deep reinforcement learning from human preferences". <i>Proceedings of the 31st International Conference on Neural Information Processing Systems</i>. NIPS'17. Red Hook, NY, USA: Curran Associates Inc. pp.&nbsp;<span class="nowrap">4302–</span>4310. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-5108-6096-4</bdi>.</cite></span>
</li>
<li id="cite_note-LessToxic-31"><span class="mw-cite-backlink">^ <a href="#cite_ref-LessToxic_31-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-LessToxic_31-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-LessToxic_31-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-LessToxic_31-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-LessToxic_31-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-LessToxic_31-5"><sup><i><b>f</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFHeaven2022" class="citation web cs1">Heaven, Will Douglas (January 27, 2022). <a rel="nofollow" class="external text" href="https://www.technologyreview.com/2022/01/27/1044398/new-gpt3-openai-chatbot-language-model-ai-toxic-misinformation/">"The new version of GPT-3 is much better behaved (and should be less toxic)"</a>. <i>MIT Technology Review</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114056/https://www.technologyreview.com/2022/01/27/1044398/new-gpt3-openai-chatbot-language-model-ai-toxic-misinformation/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-32"><span class="mw-cite-backlink"><b><a href="#cite_ref-32">^</a></b></span> <span class="reference-text"><cite id="CITEREFMohseniWangYuXiao2022" class="citation arxiv cs1">Mohseni, Sina; Wang, Haotao; Yu, Zhiding; Xiao, Chaowei; Wang, Zhangyang; Yadawa, Jay (March 7, 2022). "Taxonomy of Machine Learning Safety: A Survey and Primer". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2106.04823">2106.04823</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-33"><span class="mw-cite-backlink"><b><a href="#cite_ref-33">^</a></b></span> <span class="reference-text"><cite id="CITEREFClifton2020" class="citation web cs1">Clifton, Jesse (2020). <a rel="nofollow" class="external text" href="https://longtermrisk.org/research-agenda/">"Cooperation, Conflict, and Transformative Artificial Intelligence: A Research Agenda"</a>. <i>Center on Long-Term Risk</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230101041759/https://longtermrisk.org/research-agenda">Archived</a> from the original on January 1, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite>
<ul><li><cite id="CITEREFDafoeBachrachHadfieldHorvitz2021" class="citation journal cs1">Dafoe, Allan; Bachrach, Yoram; Hadfield, Gillian; Horvitz, Eric; Larson, Kate; Graepel, Thore (May 6, 2021). <a rel="nofollow" class="external text" href="http://www.nature.com/articles/d41586-021-01170-0">"Cooperative AI: machines must learn to find common ground"</a>. <i>Nature</i>. <b>593</b> (7857): <span class="nowrap">33–</span>36. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2021Natur.593...33D">2021Natur.593...33D</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1038%2Fd41586-021-01170-0">10.1038/d41586-021-01170-0</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0028-0836">0028-0836</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/33947992">33947992</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:233740521">233740521</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221218210857/https://www.nature.com/articles/d41586-021-01170-0">Archived</a> from the original on December 18, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-34"><span class="mw-cite-backlink"><b><a href="#cite_ref-34">^</a></b></span> <span class="reference-text"><cite id="CITEREFPrunklWhittlestone2020" class="citation book cs1">Prunkl, Carina; Whittlestone, Jess (February 7, 2020). <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/10.1145/3375627.3375803">"Beyond Near- and Long-Term"</a>. <i>Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society</i>. New York NY USA: ACM. pp.&nbsp;<span class="nowrap">138–</span>143. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1145%2F3375627.3375803">10.1145/3375627.3375803</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-4503-7110-0</bdi>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:210164673">210164673</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221016123733/https://dl.acm.org/doi/10.1145/3375627.3375803">Archived</a> from the original on October 16, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:4-35"><span class="mw-cite-backlink">^ <a href="#cite_ref-:4_35-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:4_35-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-:4_35-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-:4_35-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFIrvingAskell2019" class="citation journal cs1">Irving, Geoffrey; Askell, Amanda (February 19, 2019). <a rel="nofollow" class="external text" href="https://distill.pub/2019/safety-needs-social-scientists">"AI Safety Needs Social Scientists"</a>. <i>Distill</i>. <b>4</b> (2): 10.23915/distill.00014. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.23915%2Fdistill.00014">10.23915/distill.00014</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2476-0757">2476-0757</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:159180422">159180422</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114220/https://distill.pub/2019/safety-needs-social-scientists/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-36"><span class="mw-cite-backlink"><b><a href="#cite_ref-36">^</a></b></span> <span class="reference-text"><cite id="CITEREFGazosKahnKuscheBüscher2025" class="citation journal cs1">Gazos, Alexandros; Kahn, James; Kusche, Isabel; Büscher, Christian; Götz, Markus (April 1, 2025). <a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.ssci.2024.106731">"Organising AI for safety: Identifying structural vulnerabilities to guide the design of AI-enhanced socio-technical systems"</a>. <i>Safety Science</i>. <b>184</b> 106731. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.ssci.2024.106731">10.1016/j.ssci.2024.106731</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0925-7535">0925-7535</a>.</cite></span>
</li>
<li id="cite_note-40"><span class="mw-cite-backlink"><b><a href="#cite_ref-40">^</a></b></span> <span class="reference-text">Bringsjord, Selmer and Govindarajulu, Naveen Sundar, <a rel="nofollow" class="external text" href="https://plato.stanford.edu/archives/sum2020/entries/artificial-intelligence/">"Artificial Intelligence"</a>, The Stanford Encyclopedia of Philosophy (Summer 2020 Edition), Edward N. Zalta (ed.)</span>
</li>
<li id="cite_note-quanta_alphazero-41"><span class="mw-cite-backlink"><b><a href="#cite_ref-quanta_alphazero_41-0">^</a></b></span> <span class="reference-text"><cite class="citation news cs1"><a rel="nofollow" class="external text" href="https://www.quantamagazine.org/why-alphazeros-artificial-intelligence-has-trouble-with-the-real-world-20180221/">"Why AlphaZero's Artificial Intelligence Has Trouble With the Real World"</a>. <i>Quanta Magazine</i>. 2018<span class="reference-accessdate">. Retrieved <span class="nowrap">June 20,</span> 2020</span>.</cite></span>
</li>
<li id="cite_note-quanta_problem-42"><span class="mw-cite-backlink"><b><a href="#cite_ref-quanta_problem_42-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFWolchover2020" class="citation news cs1">Wolchover, Natalie (January 30, 2020). <a rel="nofollow" class="external text" href="https://www.quantamagazine.org/artificial-intelligence-will-do-what-we-ask-thats-a-problem-20200130/">"Artificial Intelligence Will Do What We Ask. That's a Problem"</a>. <i>Quanta Magazine</i><span class="reference-accessdate">. Retrieved <span class="nowrap">June 21,</span> 2020</span>.</cite></span>
</li>
<li id="cite_note-43"><span class="mw-cite-backlink"><b><a href="#cite_ref-43">^</a></b></span> <span class="reference-text">Bull, Larry. "On model-based evolutionary computation." Soft Computing 3, no. 2 (1999): 76–82.</span>
</li>
<li id="cite_note-Wiener1960-44"><span class="mw-cite-backlink">^ <a href="#cite_ref-Wiener1960_44-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Wiener1960_44-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWiener1960" class="citation journal cs1">Wiener, Norbert (May 6, 1960). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://www.science.org/doi/10.1126/science.131.3410.1355">"Some Moral and Technical Consequences of Automation: As machines learn they may develop unforeseen strategies at rates that baffle their programmers"</a></span>. <i>Science</i>. <b>131</b> (3410): <span class="nowrap">1355–</span>1358. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1126%2Fscience.131.3410.1355">10.1126/science.131.3410.1355</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0036-8075">0036-8075</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/17841602">17841602</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:30855376">30855376</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221015105034/https://www.science.org/doi/10.1126/science.131.3410.1355">Archived</a> from the original on October 15, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-Gabriel2020-45"><span class="mw-cite-backlink">^ <a href="#cite_ref-Gabriel2020_45-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Gabriel2020_45-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Gabriel2020_45-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Gabriel2020_45-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFGabriel2020" class="citation journal cs1">Gabriel, Iason (September 1, 2020). <a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11023-020-09539-2">"Artificial Intelligence, Values, and Alignment"</a>. <i>Minds and Machines</i>. <b>30</b> (3): <span class="nowrap">411–</span>437. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2001.09768">2001.09768</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11023-020-09539-2">10.1007/s11023-020-09539-2</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1572-8641">1572-8641</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:210920551">210920551</a>.</cite></span>
</li>
<li id="cite_note-46"><span class="mw-cite-backlink"><b><a href="#cite_ref-46">^</a></b></span> <span class="reference-text"><cite id="CITEREFThe_Ezra_Klein_Show2021" class="citation news cs1">The Ezra Klein Show (June 4, 2021). <a rel="nofollow" class="external text" href="https://www.nytimes.com/2021/06/04/opinion/ezra-klein-podcast-brian-christian.html">"If 'All Models Are Wrong,' Why Do We Give Them So Much Power?"</a>. <i>The New York Times</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0362-4331">0362-4331</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230215224050/https://www.nytimes.com/2021/06/04/opinion/ezra-klein-podcast-brian-christian.html">Archived</a> from the original on February 15, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">March 13,</span> 2023</span>.</cite>
<ul><li><cite id="CITEREFWolchover2015" class="citation web cs1">Wolchover, Natalie (April 21, 2015). <a rel="nofollow" class="external text" href="https://www.quantamagazine.org/artificial-intelligence-aligned-with-human-values-qa-with-stuart-russell-20150421/">"Concerns of an Artificial Intelligence Pioneer"</a>. <i>Quanta Magazine</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.quantamagazine.org/artificial-intelligence-aligned-with-human-values-qa-with-stuart-russell-20150421/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">March 13,</span> 2023</span>.</cite></li>
<li><cite id="CITEREFCalifornia_Assembly" class="citation web cs1">California Assembly. <a rel="nofollow" class="external text" href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201720180ACR215">"Bill Text – ACR-215 23 Asilomar AI Principles"</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201720180ACR215">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-MasteringLanguage-47"><span class="mw-cite-backlink">^ <a href="#cite_ref-MasteringLanguage_47-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-MasteringLanguage_47-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFJohnsonIziev2022" class="citation news cs1">Johnson, Steven; Iziev, Nikita (April 15, 2022). <a rel="nofollow" class="external text" href="https://www.nytimes.com/2022/04/15/magazine/ai-language.html">"A.I. Is Mastering Language. Should We Trust What It Says?"</a>. <i>The New York Times</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0362-4331">0362-4331</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221124151408/https://www.nytimes.com/2022/04/15/magazine/ai-language.html">Archived</a> from the original on November 24, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 18,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-48"><span class="mw-cite-backlink"><b><a href="#cite_ref-48">^</a></b></span> <span class="reference-text"><cite id="CITEREFOpenAI" class="citation web cs1">OpenAI. <a rel="nofollow" class="external text" href="https://openai.com/blog/our-approach-to-alignment-research">"Developing safe &amp; responsible AI"</a><span class="reference-accessdate">. Retrieved <span class="nowrap">March 13,</span> 2023</span>.</cite>
<ul><li><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://deepmindsafetyresearch.medium.com">"DeepMind Safety Research"</a>. <i>Medium</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114142/https://deepmindsafetyresearch.medium.com/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">March 13,</span> 2023</span>.</cite></li></ul>
</span></li>
<li id="cite_note-Unsolved2022-49"><span class="mw-cite-backlink">^ <a href="#cite_ref-Unsolved2022_49-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Unsolved2022_49-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Unsolved2022_49-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Unsolved2022_49-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Unsolved2022_49-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-Unsolved2022_49-5"><sup><i><b>f</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFHendrycksCarliniSchulmanSteinhardt2022" class="citation arxiv cs1">Hendrycks, Dan; <a href="Nicholas_Carlini" title="Nicholas Carlini">Carlini, Nicholas</a>; Schulman, John; Steinhardt, Jacob (June 16, 2022). "Unsolved Problems in ML Safety". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2109.13916">2109.13916</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-50"><span class="mw-cite-backlink"><b><a href="#cite_ref-50">^</a></b></span> <span class="reference-text"><cite id="CITEREFRussellNorvig2022" class="citation book cs1">Russell, Stuart J.; Norvig, Peter (2022). <a rel="nofollow" class="external text" href="https://www.pearson.com/us/higher-education/program/Russell-Artificial-Intelligence-A-Modern-Approach-4th-Edition/PGM1263338.html"><i>Artificial intelligence: a modern approach</i></a> (4th&nbsp;ed.). Pearson. pp.&nbsp;<span class="nowrap">4–</span>5. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-292-40113-3</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/1303900751">1303900751</a>.</cite></span>
</li>
<li id="cite_note-SpecGaming2020-51"><span class="mw-cite-backlink">^ <a href="#cite_ref-SpecGaming2020_51-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-SpecGaming2020_51-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-SpecGaming2020_51-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-SpecGaming2020_51-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-SpecGaming2020_51-4"><sup><i><b>e</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFKrakovnaUesatoMikulikRahtz2020" class="citation web cs1">Krakovna, Victoria; Uesato, Jonathan; Mikulik, Vladimir; Rahtz, Matthew; Everitt, Tom; Kumar, Ramana; Kenton, Zac; Leike, Jan; Legg, Shane (April 21, 2020). <a rel="nofollow" class="external text" href="https://www.deepmind.com/blog/specification-gaming-the-flip-side-of-ai-ingenuity">"Specification gaming: the flip side of AI ingenuity"</a>. <i>Deepmind</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114143/https://www.deepmind.com/blog/specification-gaming-the-flip-side-of-ai-ingenuity">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:111-52"><span class="mw-cite-backlink"><b><a href="#cite_ref-:111_52-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFManheimGarrabrant2018" class="citation arxiv cs1">Manheim, David; Garrabrant, Scott (2018). "Categorizing Variants of Goodhart's Law". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1803.04585">1803.04585</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-lfhp2017-53"><span class="mw-cite-backlink">^ <a href="#cite_ref-lfhp2017_53-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-lfhp2017_53-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-lfhp2017_53-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFAmodeiChristianoRay2017" class="citation web cs1">Amodei, Dario; Christiano, Paul; Ray, Alex (June 13, 2017). <a rel="nofollow" class="external text" href="https://openai.com/blog/deep-reinforcement-learning-from-human-preferences/">"Learning from Human Preferences"</a>. <i>OpenAI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20210103215933/https://openai.com/blog/deep-reinforcement-learning-from-human-preferences/">Archived</a> from the original on January 3, 2021<span class="reference-accessdate">. Retrieved <span class="nowrap">July 21,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-54"><span class="mw-cite-backlink"><b><a href="#cite_ref-54">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml">"Specification gaming examples in AI - master list - Google Drive"</a>. <i>docs.google.com</i>.</cite></span>
</li>
<li id="cite_note-55"><span class="mw-cite-backlink"><b><a href="#cite_ref-55">^</a></b></span> <span class="reference-text"><cite id="CITEREFClarkAmodei2016" class="citation web cs1">Clark, Jack; Amodei, Dario (December 21, 2016). <a rel="nofollow" class="external text" href="https://openai.com/research/faulty-reward-functions">"Faulty reward functions in the wild"</a>. <i>openai.com</i><span class="reference-accessdate">. Retrieved <span class="nowrap">December 30,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-TruthfulQA-56"><span class="mw-cite-backlink">^ <a href="#cite_ref-TruthfulQA_56-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-TruthfulQA_56-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-TruthfulQA_56-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLinHiltonEvans2022" class="citation journal cs1">Lin, Stephanie; Hilton, Jacob; Evans, Owain (2022). <a rel="nofollow" class="external text" href="https://aclanthology.org/2022.acl-long.229">"TruthfulQA: Measuring How Models Mimic Human Falsehoods"</a>. <i>Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</i>. Dublin, Ireland: Association for Computational Linguistics: <span class="nowrap">3214–</span>3252. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2109.07958">2109.07958</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.18653%2Fv1%2F2022.acl-long.229">10.18653/v1/2022.acl-long.229</a></span>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:237532606">237532606</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114231/https://aclanthology.org/2022.acl-long.229/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-Naughton2021-57"><span class="mw-cite-backlink">^ <a href="#cite_ref-Naughton2021_57-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Naughton2021_57-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Naughton2021_57-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFNaughton2021" class="citation news cs1">Naughton, John (October 2, 2021). <a rel="nofollow" class="external text" href="https://www.theguardian.com/commentisfree/2021/oct/02/the-truth-about-artificial-intelligence-it-isnt-that-honest">"The truth about artificial intelligence? It isn't that honest"</a>. <i>The Observer</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0029-7712">0029-7712</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230213231317/https://www.theguardian.com/commentisfree/2021/oct/02/the-truth-about-artificial-intelligence-it-isnt-that-honest">Archived</a> from the original on February 13, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-58"><span class="mw-cite-backlink"><b><a href="#cite_ref-58">^</a></b></span> <span class="reference-text"><cite id="CITEREFJiLeeFrieskeYu2022" class="citation journal cs1">Ji, Ziwei; Lee, Nayeon; Frieske, Rita; Yu, Tiezheng; Su, Dan; Xu, Yan; Ishii, Etsuko; Bang, Yejin; Madotto, Andrea; Fung, Pascale (February 1, 2022). "Survey of Hallucination in Natural Language Generation". <i>ACM Computing Surveys</i>. <b>55</b> (12): <span class="nowrap">1–</span>38. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2202.03629">2202.03629</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1145%2F3571730">10.1145/3571730</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:246652372">246652372</a>.</cite>
<ul><li><cite id="CITEREFElse2023" class="citation journal cs1">Else, Holly (January 12, 2023). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://www.nature.com/articles/d41586-023-00056-7">"Abstracts written by ChatGPT fool scientists"</a></span>. <i>Nature</i>. <b>613</b> (7944): 423. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2023Natur.613..423E">2023Natur.613..423E</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1038%2Fd41586-023-00056-7">10.1038/d41586-023-00056-7</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/36635510">36635510</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:255773668">255773668</a>.</cite></li></ul>
</span></li>
<li id="cite_note-:5-59"><span class="mw-cite-backlink"><b><a href="#cite_ref-:5_59-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFRussell" class="citation web cs1">Russell, Stuart. <a rel="nofollow" class="external text" href="https://www.edge.org/conversation/the-myth-of-ai">"Of Myths and Moonshine"</a>. <i>Edge.org</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.edge.org/conversation/the-myth-of-ai">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 19,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-60"><span class="mw-cite-backlink"><b><a href="#cite_ref-60">^</a></b></span> <span class="reference-text"><cite id="CITEREFTasioulas2019" class="citation journal cs1">Tasioulas, John (2019). "First Steps Towards an Ethics of Robots and Artificial Intelligence". <i>Journal of Practical Ethics</i>. <b>7</b> (1): <span class="nowrap">61–</span>95.</cite></span>
</li>
<li id="cite_note-61"><span class="mw-cite-backlink"><b><a href="#cite_ref-61">^</a></b></span> <span class="reference-text"><cite id="CITEREFBooth2025" class="citation magazine cs1">Booth, Harry (February 19, 2025). <a rel="nofollow" class="external text" href="https://time.com/7259395/ai-chess-cheating-palisade-research/">"When AI Thinks It Will Lose, It Sometimes Cheats"</a>. <i>TIME</i><span class="reference-accessdate">. Retrieved <span class="nowrap">February 23,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-:722-62"><span class="mw-cite-backlink"><b><a href="#cite_ref-:722_62-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFWellsDeepa_SeetharamanHorwitz2021" class="citation news cs1">Wells, Georgia; Deepa Seetharaman; Horwitz, Jeff (November 5, 2021). <a rel="nofollow" class="external text" href="https://www.wsj.com/articles/facebook-bad-for-you-360-million-users-say-yes-company-documents-facebook-files-11636124681">"Is Facebook Bad for You? It Is for About 360 Million Users, Company Surveys Suggest"</a>. <i>The Wall Street Journal</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0099-9660">0099-9660</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.wsj.com/articles/facebook-bad-for-you-360-million-users-say-yes-company-documents-facebook-files-11636124681">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 19,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:822-63"><span class="mw-cite-backlink"><b><a href="#cite_ref-:822_63-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFBarrettHendrixSims2021" class="citation report cs1">Barrett, Paul M.; Hendrix, Justin; Sims, J. Grant (September 2021). <a rel="nofollow" class="external text" href="https://bhr.stern.nyu.edu/polarization-report-page">How Social Media Intensifies U.S. Political Polarization-And What Can Be Done About It</a> (Report). Center for Business and Human Rights, NYU. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230201180005/https://bhr.stern.nyu.edu/polarization-report-page">Archived</a> from the original on February 1, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-64"><span class="mw-cite-backlink"><b><a href="#cite_ref-64">^</a></b></span> <span class="reference-text"><cite id="CITEREFShepardson2018" class="citation news cs1">Shepardson, David (May 24, 2018). <a rel="nofollow" class="external text" href="https://www.reuters.com/article/us-uber-crash-idUSKCN1IP26K">"Uber disabled emergency braking in self-driving car: U.S. agency"</a>. <i>Reuters</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.reuters.com/article/us-uber-crash-idUSKCN1IP26K">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 20,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-65"><span class="mw-cite-backlink"><b><a href="#cite_ref-65">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.technologyreview.com/2020/02/17/844721/ai-openai-moonshot-elon-musk-sam-altman-greg-brockman-messy-secretive-reality/">"The messy, secretive reality behind OpenAI's bid to save the world"</a>. <i>MIT Technology Review</i><span class="reference-accessdate">. Retrieved <span class="nowrap">August 25,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-66"><span class="mw-cite-backlink"><b><a href="#cite_ref-66">^</a></b></span> <span class="reference-text"><cite id="CITEREFHeath2024" class="citation web cs1">Heath, Alex (January 18, 2024). <a rel="nofollow" class="external text" href="https://www.theverge.com/2024/1/18/24042354/mark-zuckerberg-meta-agi-reorg-interview">"Mark Zuckerberg's new goal is creating artificial general intelligence"</a>. <i>The Verge</i><span class="reference-accessdate">. Retrieved <span class="nowrap">November 5,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-67"><span class="mw-cite-backlink"><b><a href="#cite_ref-67">^</a></b></span> <span class="reference-text"><cite id="CITEREFJohnson" class="citation web cs1">Johnson, Dave. <a rel="nofollow" class="external text" href="https://www.businessinsider.com/google-deepmind">"DeepMind is Google's AI research hub. Here's what it does, where it's located, and how it differs from OpenAI"</a>. <i>Business Insider</i><span class="reference-accessdate">. Retrieved <span class="nowrap">August 25,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-eallm2022-68"><span class="mw-cite-backlink">^ <a href="#cite_ref-eallm2022_68-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-eallm2022_68-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWeiTayBommasaniRaffel2022" class="citation journal cs1">Wei, Jason; Tay, Yi; Bommasani, Rishi; Raffel, Colin; Zoph, Barret; Borgeaud, Sebastian; Yogatama, Dani; Bosma, Maarten; Zhou, Denny; Metzler, Donald; Chi, Ed H.; Hashimoto, Tatsunori; Vinyals, Oriol; Liang, Percy; Dean, Jeff; Fedus, William (October 26, 2022). "Emergent Abilities of Large Language Models". <i>Transactions on Machine Learning Research</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2206.07682">2206.07682</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2835-8856">2835-8856</a>.</cite></span>
</li>
<li id="cite_note-:0-69"><span class="mw-cite-backlink">^ <a href="#cite_ref-:0_69-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:0_69-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFCaballeroGuptaRishKrueger2022" class="citation arxiv cs1">Caballero, Ethan; Gupta, Kshitij; Rish, Irina; Krueger, David (2022). "Broken Neural Scaling Laws". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.14891">2210.14891</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-70"><span class="mw-cite-backlink"><b><a href="#cite_ref-70">^</a></b></span> <span class="reference-text"><cite id="CITEREFDominguez2022" class="citation web cs1">Dominguez, Daniel (May 19, 2022). <a rel="nofollow" class="external text" href="https://www.infoq.com/news/2022/05/deepmind-gato-ai-agent/">"DeepMind Introduces Gato, a New Generalist AI Agent"</a>. <i>InfoQ</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.infoq.com/news/2022/05/deepmind-gato-ai-agent/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 9,</span> 2022</span>.</cite>
<ul><li><cite id="CITEREFEdwards2022" class="citation web cs1">Edwards, Ben (April 26, 2022). <a rel="nofollow" class="external text" href="https://arstechnica.com/information-technology/2022/09/new-ai-assistant-can-browse-search-and-use-web-apps-like-a-human/">"Adept's AI assistant can browse, search, and use web apps like a human"</a>. <i>Ars Technica</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230117194921/https://arstechnica.com/information-technology/2022/09/new-ai-assistant-can-browse-search-and-use-web-apps-like-a-human/">Archived</a> from the original on January 17, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 9,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-71"><span class="mw-cite-backlink"><b><a href="#cite_ref-71">^</a></b></span> <span class="reference-text"><cite id="CITEREFGraceStewartSandkühlerThomas2024" class="citation arxiv cs1">Grace, Katja; Stewart, Harlan; Sandkühler, Julia Fabienne; Thomas, Stephen; Weinstein-Raun, Ben; Brauner, Jan (January 5, 2024). "Thousands of AI Authors on the Future of AI". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2401.02843">2401.02843</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-:2822-72"><span class="mw-cite-backlink"><b><a href="#cite_ref-:2822_72-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFGraceSalvatierDafoeZhang2018" class="citation journal cs1">Grace, Katja; Salvatier, John; Dafoe, Allan; Zhang, Baobao; Evans, Owain (July 31, 2018). <a rel="nofollow" class="external text" href="http://jair.org/index.php/jair/article/view/11222">"Viewpoint: When Will AI Exceed Human Performance? Evidence from AI Experts"</a>. <i>Journal of Artificial Intelligence Research</i>. <b>62</b>: <span class="nowrap">729–</span>754. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1613%2Fjair.1.11222">10.1613/jair.1.11222</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1076-9757">1076-9757</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:8746462">8746462</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114220/https://jair.org/index.php/jair/article/view/11222">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:2922-73"><span class="mw-cite-backlink"><b><a href="#cite_ref-:2922_73-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFZhangAnderljungKahnDreksler2021" class="citation journal cs1">Zhang, Baobao; Anderljung, Markus; Kahn, Lauren; Dreksler, Noemi; Horowitz, Michael C.; Dafoe, Allan (August 2, 2021). <a rel="nofollow" class="external text" href="https://jair.org/index.php/jair/article/view/12895">"Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers"</a>. <i>Journal of Artificial Intelligence Research</i>. <b>71</b>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2105.02117">2105.02117</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1613%2Fjair.1.12895">10.1613/jair.1.12895</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1076-9757">1076-9757</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:233740003">233740003</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114143/https://jair.org/index.php/jair/article/view/12895">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:1701-74"><span class="mw-cite-backlink"><b><a href="#cite_ref-:1701_74-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFFuture_of_Life_Institute2023" class="citation web cs1">Future of Life Institute (March 22, 2023). <a rel="nofollow" class="external text" href="https://futureoflife.org/open-letter/pause-giant-ai-experiments/">"Pause Giant AI Experiments: An Open Letter"</a><span class="reference-accessdate">. Retrieved <span class="nowrap">April 20,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-75"><span class="mw-cite-backlink"><b><a href="#cite_ref-75">^</a></b></span> <span class="reference-text"><cite id="CITEREFWangMaFengZhang2024" class="citation journal cs1">Wang, Lei; Ma, Chen; Feng, Xueyang; Zhang, Zeyu; Yang, Hao; Zhang, Jingsen; Chen, Zhiyuan; Tang, Jiakai; Chen, Xu (2024). <a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2023arXiv230811432W">"A survey on large language model based autonomous agents"</a>. <i>Frontiers of Computer Science</i>. <b>18</b> (6) 186345. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2308.11432">2308.11432</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11704-024-40231-1">10.1007/s11704-024-40231-1</a><span class="reference-accessdate">. Retrieved <span class="nowrap">February 11,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-76"><span class="mw-cite-backlink"><b><a href="#cite_ref-76">^</a></b></span> <span class="reference-text"><cite id="CITEREFBerglundSticklandBalesniKaufmann2023" class="citation arxiv cs1">Berglund, Lukas; Stickland, Asa Cooper; Balesni, Mikita; Kaufmann, Max; Tong, Meg; Korbak, Tomasz; Kokotajlo, Daniel; Evans, Owain (September 1, 2023). "Taken out of context: On measuring situational awareness in LLMs". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2309.00667">2309.00667</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-77"><span class="mw-cite-backlink"><b><a href="#cite_ref-77">^</a></b></span> <span class="reference-text"><cite id="CITEREFLaineMeinkeEvans2023" class="citation journal cs1">Laine, Rudolf; Meinke, Alexander; Evans, Owain (November 28, 2023). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=DRk4bWKr41&amp;referrer=%5Bthe+profile+of+Rudolf+Laine%5D(/profile?id=~Rudolf_Laine1)">"Towards a Situational Awareness Benchmark for LLMs"</a>. <i>NeurIPS 2023 SoLaR Workshop</i>.</cite></span>
</li>
<li id="cite_note-:3-78"><span class="mw-cite-backlink">^ <a href="#cite_ref-:3_78-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:3_78-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPanShernZouLi2023" class="citation journal cs1">Pan, Alexander; Shern, Chan Jun; Zou, Andy; Li, Nathaniel; Basart, Steven; Woodside, Thomas; Ng, Jonathan; Zhang, Emmons; Scott, Dan; Hendrycks (April 3, 2023). "Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark". <i>Proceedings of the 40th International Conference on Machine Learning</i>. PMLR. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2304.03279">2304.03279</a></span>.</cite></span>
</li>
<li id="cite_note-dllmmwe2022-79"><span class="mw-cite-backlink">^ <a href="#cite_ref-dllmmwe2022_79-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-dllmmwe2022_79-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-dllmmwe2022_79-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-dllmmwe2022_79-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPerezRingerLukošiūtėNguyen2022" class="citation arxiv cs1">Perez, Ethan; Ringer, Sam; Lukošiūtė, Kamilė; Nguyen, Karina; Chen, Edwin; Heiner, Scott; Pettit, Craig; Olsson, Catherine; Kundu, Sandipan; Kadavath, Saurav; Jones, Andy; Chen, Anna; Mann, Ben; Israel, Brian; Seethor, Bryan (December 19, 2022). "Discovering Language Model Behaviors with Model-Written Evaluations". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2212.09251">2212.09251</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-80"><span class="mw-cite-backlink"><b><a href="#cite_ref-80">^</a></b></span> <span class="reference-text"><cite id="CITEREFOrseauArmstrong2016" class="citation journal cs1">Orseau, Laurent; Armstrong, Stuart (June 25, 2016). <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/abs/10.5555/3020948.3021006">"Safely interruptible agents"</a>. <i>Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence</i>. UAI'16. Arlington, Virginia, USA: AUAI Press: <span class="nowrap">557–</span>566. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-9966431-1-5</bdi>.</cite></span>
</li>
<li id="cite_note-Gridworlds-81"><span class="mw-cite-backlink">^ <a href="#cite_ref-Gridworlds_81-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Gridworlds_81-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLeikeMarticKrakovnaOrtega2017" class="citation arxiv cs1">Leike, Jan; Martic, Miljan; Krakovna, Victoria; Ortega, Pedro A.; Everitt, Tom; Lefrancq, Andrew; Orseau, Laurent; Legg, Shane (November 28, 2017). "AI Safety Gridworlds". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1711.09883">1711.09883</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-OffSwitch-82"><span class="mw-cite-backlink">^ <a href="#cite_ref-OffSwitch_82-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-OffSwitch_82-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-OffSwitch_82-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-OffSwitch_82-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFHadfield-MenellDraganAbbeelRussell2017" class="citation journal cs1">Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (August 19, 2017). <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/10.5555/3171642.3171675">"The off-switch game"</a>. <i>Proceedings of the 26th International Joint Conference on Artificial Intelligence</i>. IJCAI'17. Melbourne, Australia: AAAI Press: <span class="nowrap">220–</span>227. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-9992411-0-3</bdi>.</cite></span>
</li>
<li id="cite_note-optsp-83"><span class="mw-cite-backlink">^ <a href="#cite_ref-optsp_83-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-optsp_83-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-optsp_83-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-optsp_83-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFTurnerSmithShahCritch2021" class="citation conference cs1">Turner, Alexander Matt; Smith, Logan Riggs; Shah, Rohin; Critch, Andrew; Tadepalli, Prasad (2021). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=l7-DBWawSZH">"Optimal policies tend to seek power"</a>. <i>Advances in neural information processing systems</i>.</cite></span>
</li>
<li id="cite_note-84"><span class="mw-cite-backlink"><b><a href="#cite_ref-84">^</a></b></span> <span class="reference-text"><cite id="CITEREFTurnerTadepalli2022" class="citation conference cs1">Turner, Alexander Matt; Tadepalli, Prasad (2022). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=GFgjnk2Q-ju">"Parametrically retargetable decision-makers tend to seek power"</a>. <i>Advances in neural information processing systems</i>.</cite></span>
</li>
<li id="cite_note-Superintelligence-85"><span class="mw-cite-backlink">^ <a href="#cite_ref-Superintelligence_85-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Superintelligence_85-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Superintelligence_85-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-Superintelligence_85-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-Superintelligence_85-4"><sup><i><b>e</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFBostrom2014" class="citation book cs1">Bostrom, Nick (2014). <i>Superintelligence: Paths, Dangers, Strategies</i> (1st&nbsp;ed.). USA: Oxford University Press, Inc. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-19-967811-2</bdi>.</cite></span>
</li>
<li id="cite_note-:1-86"><span class="mw-cite-backlink">^ <a href="#cite_ref-:1_86-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:1_86-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.safe.ai/statement-on-ai-risk">"Statement on AI Risk | CAIS"</a>. <i>www.safe.ai</i><span class="reference-accessdate">. Retrieved <span class="nowrap">July 17,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-87"><span class="mw-cite-backlink"><b><a href="#cite_ref-87">^</a></b></span> <span class="reference-text"><cite id="CITEREFRoose2023" class="citation news cs1">Roose, Kevin (May 30, 2023). <a rel="nofollow" class="external text" href="https://www.nytimes.com/2023/05/30/technology/ai-threat-warning.html">"A.I. Poses 'Risk of Extinction,' Industry Leaders Warn"</a>. <i>The New York Times</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0362-4331">0362-4331</a><span class="reference-accessdate">. Retrieved <span class="nowrap">July 17,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-88"><span class="mw-cite-backlink"><b><a href="#cite_ref-88">^</a></b></span> <span class="reference-text"><cite id="CITEREFTuring1951" class="citation speech cs1">Turing, Alan (1951). <a rel="nofollow" class="external text" href="https://turingarchive.kings.cam.ac.uk/publications-lectures-and-talks-amtb/amt-b-4"><i>Intelligent machinery, a heretical theory</i></a> (Speech). Lecture given to '51 Society'. Manchester: The Turing Digital Archive. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220926004549/https://turingarchive.kings.cam.ac.uk/publications-lectures-and-talks-amtb/amt-b-4">Archived</a> from the original on September 26, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 22,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-89"><span class="mw-cite-backlink"><b><a href="#cite_ref-89">^</a></b></span> <span class="reference-text"><cite id="CITEREFTuring1951" class="citation episode cs1">Turing, Alan (May 15, 1951). "Can digital computers think?". <i>Automatic Calculating Machines</i>. Episode 2. BBC. <a rel="nofollow" class="external text" href="https://turingarchive.kings.cam.ac.uk/publications-lectures-and-talks-amtb/amt-b-6">Can digital computers think?</a>.</cite></span>
</li>
<li id="cite_note-:3022-91"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3022_91-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFMuehlhauser2016" class="citation web cs1">Muehlhauser, Luke (January 29, 2016). <a rel="nofollow" class="external text" href="https://lukemuehlhauser.com/sutskever-on-talking-machines/">"Sutskever on Talking Machines"</a>. <i>Luke Muehlhauser</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220927200137/https://lukemuehlhauser.com/sutskever-on-talking-machines/">Archived</a> from the original on September 27, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:3122-93"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3122_93-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFShanahan2015" class="citation book cs1">Shanahan, Murray (2015). <i>The technological singularity</i>. Cambridge, Massachusetts: MIT Press. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-262-52780-4</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/917889148">917889148</a>.</cite></span>
</li>
<li id="cite_note-:3322-95"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3322_95-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFRossi" class="citation news cs1">Rossi, Francesca. <a rel="nofollow" class="external text" href="https://www.washingtonpost.com/news/in-theory/wp/2015/11/05/how-do-you-teach-a-machine-to-be-moral/">"How do you teach a machine to be moral?"</a>. <i>The Washington Post</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0190-8286">0190-8286</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.washingtonpost.com/news/in-theory/wp/2015/11/05/how-do-you-teach-a-machine-to-be-moral/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:3422-96"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3422_96-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFAaronson2022" class="citation web cs1">Aaronson, Scott (June 17, 2022). <a rel="nofollow" class="external text" href="https://scottaaronson.blog/?p=6484">"OpenAI!"</a>. <i>Shtetl-Optimized</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220827214238/https://scottaaronson.blog/?p=6484">Archived</a> from the original on August 27, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:3522-97"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3522_97-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFSelman" class="citation cs2">Selman, Bart, <a rel="nofollow" class="external text" href="https://futureoflife.org/data/PDF/bart_selman.pdf"><i>Intelligence Explosion: Science or Fiction?</i></a> <span class="cs1-format">(PDF)</span>, <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220531022540/https://futureoflife.org/data/PDF/bart_selman.pdf">archived</a> <span class="cs1-format">(PDF)</span> from the original on May 31, 2022<span class="reference-accessdate">, retrieved <span class="nowrap">September 12,</span> 2022</span></cite></span>
</li>
<li id="cite_note-:3622-98"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3622_98-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFMcAllester2014" class="citation web cs1">McAllester (August 10, 2014). <a rel="nofollow" class="external text" href="https://machinethoughts.wordpress.com/2014/08/10/friendly-ai-and-the-servant-mission/">"Friendly AI and the Servant Mission"</a>. <i>Machine Thoughts</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220928054922/https://machinethoughts.wordpress.com/2014/08/10/friendly-ai-and-the-servant-mission/">Archived</a> from the original on September 28, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-AGISafetyLitReview-99"><span class="mw-cite-backlink">^ <a href="#cite_ref-AGISafetyLitReview_99-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-AGISafetyLitReview_99-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-AGISafetyLitReview_99-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-AGISafetyLitReview_99-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-AGISafetyLitReview_99-4"><sup><i><b>e</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFEverittLeaHutter2018" class="citation arxiv cs1">Everitt, Tom; Lea, Gary; Hutter, Marcus (May 21, 2018). "AGI Safety Literature Review". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1805.01109">1805.01109</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-:3822-100"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3822_100-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFShane2009" class="citation web cs1">Shane (August 31, 2009). <a rel="nofollow" class="external text" href="http://www.vetta.org/2009/08/funding-safe-agi/">"Funding safe AGI"</a>. <i>vetta project</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221010143110/http://www.vetta.org/2009/08/funding-safe-agi/">Archived</a> from the original on October 10, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:3922-101"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3922_101-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFHorvitz2016" class="citation web cs1">Horvitz, Eric (June 27, 2016). <a rel="nofollow" class="external text" href="http://erichorvitz.com/OSTP-CMU_AI_Safety_framing_talk.pdf">"Reflections on Safety and Artificial Intelligence"</a> <span class="cs1-format">(PDF)</span>. <i>Eric Horvitz</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221010143106/http://erichorvitz.com/OSTP-CMU_AI_Safety_framing_talk.pdf">Archived</a> <span class="cs1-format">(PDF)</span> from the original on October 10, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">April 20,</span> 2020</span>.</cite></span>
</li>
<li id="cite_note-:4022-102"><span class="mw-cite-backlink"><b><a href="#cite_ref-:4022_102-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFChollet2018" class="citation web cs1">Chollet, François (December 8, 2018). <a rel="nofollow" class="external text" href="https://medium.com/@francois.chollet/the-impossibility-of-intelligence-explosion-5be4a9eda6ec">"The implausibility of intelligence explosion"</a>. <i>Medium</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20210322214203/https://medium.com/@francois.chollet/the-impossibility-of-intelligence-explosion-5be4a9eda6ec">Archived</a> from the original on March 22, 2021<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:4122-103"><span class="mw-cite-backlink"><b><a href="#cite_ref-:4122_103-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFMarcus2022" class="citation web cs1">Marcus, Gary (June 6, 2022). <a rel="nofollow" class="external text" href="https://www.scientificamerican.com/article/artificial-general-intelligence-is-not-as-imminent-as-you-might-think1/">"Artificial General Intelligence Is Not as Imminent as You Might Think"</a>. <i>Scientific American</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220915154158/https://www.scientificamerican.com/article/artificial-general-intelligence-is-not-as-imminent-as-you-might-think1/">Archived</a> from the original on September 15, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-:4322-104"><span class="mw-cite-backlink"><b><a href="#cite_ref-:4322_104-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFBarber2016" class="citation web cs1">Barber, Lynsey (July 31, 2016). <a rel="nofollow" class="external text" href="https://www.cityam.com/phew-facebooks-ai-chief-says-intelligent-machines-not/">"Phew! Facebook's AI chief says intelligent machines are not a threat to humanity"</a>. <i>CityAM</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220826063808/https://www.cityam.com/phew-facebooks-ai-chief-says-intelligent-machines-not/">Archived</a> from the original on August 26, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-105"><span class="mw-cite-backlink"><b><a href="#cite_ref-105">^</a></b></span> <span class="reference-text"><cite id="CITEREFEtzioni2016" class="citation web cs1">Etzioni, Oren (September 20, 2016). <a rel="nofollow" class="external text" href="https://www.technologyreview.com/2016/09/20/70131/no-the-experts-dont-think-superintelligent-ai-is-a-threat-to-humanity/">"No, the Experts Don't Think Superintelligent AI is a Threat to Humanity"</a>. <i>MIT Technology Review</i><span class="reference-accessdate">. Retrieved <span class="nowrap">June 10,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-106"><span class="mw-cite-backlink"><b><a href="#cite_ref-106">^</a></b></span> <span class="reference-text"><cite id="CITEREFRochonRossi2015" class="citation book cs1">Rochon, Louis-Philippe; Rossi, Sergio (February 27, 2015). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=6kzfBgAAQBAJ"><i>The Encyclopedia of Central Banking</i></a>. Edward Elgar Publishing. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-78254-744-0</bdi>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114225/https://books.google.com/books?id=6kzfBgAAQBAJ">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 13,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-107"><span class="mw-cite-backlink"><b><a href="#cite_ref-107">^</a></b></span> <span class="reference-text"><cite id="CITEREFNgRussell2000" class="citation journal cs1">Ng, Andrew Y.; Russell, Stuart J. (June 29, 2000). <a rel="nofollow" class="external text" href="https://dl.acm.org/doi/10.5555/645529.657801">"Algorithms for Inverse Reinforcement Learning"</a>. <i>Proceedings of the Seventeenth International Conference on Machine Learning</i>. ICML '00. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc.: <span class="nowrap">663–</span>670. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-55860-707-1</bdi>.</cite></span>
</li>
<li id="cite_note-108"><span class="mw-cite-backlink"><b><a href="#cite_ref-108">^</a></b></span> <span class="reference-text"><cite id="CITEREFHadfield-MenellRussellAbbeelDragan2016" class="citation conference cs1">Hadfield-Menell, Dylan; Russell, Stuart J; Abbeel, Pieter; Dragan, Anca (2016). "Cooperative inverse reinforcement learning". <i>Advances in neural information processing systems</i>. Vol.&nbsp;29. Curran Associates, Inc.</cite></span>
</li>
<li id="cite_note-109"><span class="mw-cite-backlink"><b><a href="#cite_ref-109">^</a></b></span> <span class="reference-text"><cite id="CITEREFMindermannArmstrong2018" class="citation conference cs1">Mindermann, Soren; Armstrong, Stuart (2018). "Occam's razor is insufficient to infer the preferences of irrational agents". <i>Proceedings of the 32nd international conference on neural information processing systems</i>. NIPS'18. Red Hook, NY, USA: Curran Associates Inc. pp.&nbsp;<span class="nowrap">5603–</span>5614.</cite></span>
</li>
<li id="cite_note-110"><span class="mw-cite-backlink"><b><a href="#cite_ref-110">^</a></b></span> <span class="reference-text"><cite id="CITEREFFürnkranzHüllermeierRudinSlowinski2014" class="citation journal cs1">Fürnkranz, Johannes; Hüllermeier, Eyke; Rudin, Cynthia; Slowinski, Roman; Sanner, Scott (2014). <a rel="nofollow" class="external text" href="http://drops.dagstuhl.de/opus/volltexte/2014/4550/">"Preference Learning"</a>. <i>Dagstuhl Reports</i>. <b>4</b> (3). Marc Herbstritt: 27 pages. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.4230%2FDAGREP.4.3.1">10.4230/DAGREP.4.3.1</a></span>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114221/https://drops.dagstuhl.de/opus/volltexte/2014/4550/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-111"><span class="mw-cite-backlink"><b><a href="#cite_ref-111">^</a></b></span> <span class="reference-text"><cite id="CITEREFGaoSchulmanHilton2022" class="citation arxiv cs1">Gao, Leo; Schulman, John; Hilton, Jacob (October 19, 2022). "Scaling Laws for Reward Model Overoptimization". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.10760">2210.10760</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-112"><span class="mw-cite-backlink"><b><a href="#cite_ref-112">^</a></b></span> <span class="reference-text"><cite id="CITEREFAnderson2022" class="citation web cs1">Anderson, Martin (April 5, 2022). <a rel="nofollow" class="external text" href="https://www.unite.ai/the-perils-of-using-quotations-to-authenticate-nlg-content/">"The Perils of Using Quotations to Authenticate NLG Content"</a>. <i>Unite.AI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114139/https://www.unite.ai/the-perils-of-using-quotations-to-authenticate-nlg-content/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 21,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-Wiggers2022-113"><span class="mw-cite-backlink">^ <a href="#cite_ref-Wiggers2022_113-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Wiggers2022_113-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWiggers2022" class="citation web cs1">Wiggers, Kyle (February 5, 2022). <a rel="nofollow" class="external text" href="https://venturebeat.com/2022/02/05/despite-recent-progress-ai-powered-chatbots-still-have-a-long-way-to-go/">"Despite recent progress, AI-powered chatbots still have a long way to go"</a>. <i>VentureBeat</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220723184144/https://venturebeat.com/2022/02/05/despite-recent-progress-ai-powered-chatbots-still-have-a-long-way-to-go/">Archived</a> from the original on July 23, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-114"><span class="mw-cite-backlink"><b><a href="#cite_ref-114">^</a></b></span> <span class="reference-text"><cite id="CITEREFHendrycksBurnsBasartCritch2021" class="citation journal cs1">Hendrycks, Dan; Burns, Collin; Basart, Steven; Critch, Andrew; Li, Jerry; Song, Dawn; Steinhardt, Jacob (July 24, 2021). "Aligning AI With Shared Human Values". <i>International Conference on Learning Representations</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2008.02275">2008.02275</a></span>.</cite></span>
</li>
<li id="cite_note-115"><span class="mw-cite-backlink"><b><a href="#cite_ref-115">^</a></b></span> <span class="reference-text"><cite id="CITEREFPerezHuangSongCai2022" class="citation arxiv cs1">Perez, Ethan; Huang, Saffron; Song, Francis; Cai, Trevor; Ring, Roman; Aslanides, John; Glaese, Amelia; McAleese, Nat; Irving, Geoffrey (February 7, 2022). "Red Teaming Language Models with Language Models". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2202.03286">2202.03286</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite>
<ul><li><cite id="CITEREFBhattacharyya2022" class="citation web cs1">Bhattacharyya, Sreejani (February 14, 2022). <a rel="nofollow" class="external text" href="https://analyticsindiamag.com/deepminds-red-teaming-language-models-with-language-models-what-is-it/">"DeepMind's "red teaming" language models with language models: What is it?"</a>. <i>Analytics India Magazine</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230213145212/https://analyticsindiamag.com/deepminds-red-teaming-language-models-with-language-models-what-is-it/">Archived</a> from the original on February 13, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-116"><span class="mw-cite-backlink"><b><a href="#cite_ref-116">^</a></b></span> <span class="reference-text"><cite id="CITEREFAndersonAnderson2007" class="citation journal cs1">Anderson, Michael; Anderson, Susan Leigh (December 15, 2007). <a rel="nofollow" class="external text" href="https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2065">"Machine Ethics: Creating an Ethical Intelligent Agent"</a>. <i>AI Magazine</i>. <b>28</b> (4): 15. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1609%2Faimag.v28i4.2065">10.1609/aimag.v28i4.2065</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2371-9621">2371-9621</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:17033332">17033332</a><span class="reference-accessdate">. Retrieved <span class="nowrap">March 14,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-117"><span class="mw-cite-backlink"><b><a href="#cite_ref-117">^</a></b></span> <span class="reference-text"><cite id="CITEREFWiegel2010" class="citation journal cs1">Wiegel, Vincent (December 1, 2010). <a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs10676-010-9239-1">"Wendell Wallach and Colin Allen: moral machines: teaching robots right from wrong"</a>. <i>Ethics and Information Technology</i>. <b>12</b> (4): <span class="nowrap">359–</span>361. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs10676-010-9239-1">10.1007/s10676-010-9239-1</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1572-8439">1572-8439</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:30532107">30532107</a>.</cite></span>
</li>
<li id="cite_note-118"><span class="mw-cite-backlink"><b><a href="#cite_ref-118">^</a></b></span> <span class="reference-text"><cite id="CITEREFWallachAllen2009" class="citation book cs1">Wallach, Wendell; Allen, Colin (2009). <a rel="nofollow" class="external text" href="https://oxford.universitypressscholarship.com/10.1093/acprof:oso/9780195374049.001.0001/acprof-9780195374049"><i>Moral Machines: Teaching Robots Right from Wrong</i></a>. New York: Oxford University Press. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-19-537404-9</bdi>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230315193012/https://academic.oup.com/pages/op-migration-welcome">Archived</a> from the original on March 15, 2023.</cite></span>
</li>
<li id="cite_note-Phelps2023-120"><span class="mw-cite-backlink">^ <a href="#cite_ref-Phelps2023_120-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Phelps2023_120-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPhelpsRanson2023" class="citation arxiv cs1">Phelps, Steve; Ranson, Rebecca (2023). "Of Models and Tin-Men - A Behavioral Economics Study of Principal-Agent Problems in AI Alignment Using Large-Language Models". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2307.11137">2307.11137</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-121"><span class="mw-cite-backlink"><b><a href="#cite_ref-121">^</a></b></span> <span class="reference-text"><cite id="CITEREFHendrycksBurnsBasartCritch2020" class="citation arxiv cs1">Hendrycks, Dan; Burns, Collin; Basart, Steven; Critch, Andrew; Li, Jerry; Song, Dawn; Steinhardt, Jacob (2020). "Aligning AI With Shared Human Values". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2008.02275">2008.02275</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-122"><span class="mw-cite-backlink"><b><a href="#cite_ref-122">^</a></b></span> <span class="reference-text"><cite id="CITEREFMacAskill2022" class="citation book cs1">MacAskill, William (2022). <a rel="nofollow" class="external text" href="https://whatweowethefuture.com/"><i>What we owe the future</i></a>. New York, NY: Basic Books, Hachette Book Group. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-5416-1862-6</bdi>. <a href="OCLC_(identifier)" class="mw-redirect" title="OCLC (identifier)">OCLC</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/oclc/1314633519">1314633519</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220914030758/https://www.basicbooks.com/titles/william-macaskill/what-we-owe-the-future/9781541618633/">Archived</a> from the original on September 14, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 11,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-RecursivelySummarizing-123"><span class="mw-cite-backlink">^ <a href="#cite_ref-RecursivelySummarizing_123-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-RecursivelySummarizing_123-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWuOuyangZieglerStiennon2021" class="citation arxiv cs1">Wu, Jeff; Ouyang, Long; Ziegler, Daniel M.; Stiennon, Nisan; Lowe, Ryan; Leike, Jan; Christiano, Paul (September 27, 2021). "Recursively Summarizing Books with Human Feedback". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2109.10862">2109.10862</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-124"><span class="mw-cite-backlink"><b><a href="#cite_ref-124">^</a></b></span> <span class="reference-text"><cite id="CITEREFPearceAhmadTanDolan-Gavitt2022" class="citation book cs1">Pearce, Hammond; Ahmad, Baleegh; Tan, Benjamin; Dolan-Gavitt, Brendan; Karri, Ramesh (2022). "Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions". <i>2022 IEEE Symposium on Security and Privacy (SP)</i>. San Francisco, CA, USA: IEEE. pp.&nbsp;<span class="nowrap">754–</span>768. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2108.09293">2108.09293</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FSP46214.2022.9833571">10.1109/SP46214.2022.9833571</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-6654-1316-9</bdi>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:245220588">245220588</a>.</cite></span>
</li>
<li id="cite_note-125"><span class="mw-cite-backlink"><b><a href="#cite_ref-125">^</a></b></span> <span class="reference-text"><cite id="CITEREFIrvingAmodei2018" class="citation web cs1">Irving, Geoffrey; Amodei, Dario (May 3, 2018). <a rel="nofollow" class="external text" href="https://openai.com/blog/debate/">"AI Safety via Debate"</a>. <i>OpenAI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://openai.com/blog/debate/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-sslawe-126"><span class="mw-cite-backlink">^ <a href="#cite_ref-sslawe_126-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-sslawe_126-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFChristianoShlegerisAmodei2018" class="citation arxiv cs1">Christiano, Paul; Shlegeris, Buck; Amodei, Dario (October 19, 2018). "Supervising strong learners by amplifying weak experts". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1810.08575">1810.08575</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-127"><span class="mw-cite-backlink"><b><a href="#cite_ref-127">^</a></b></span> <span class="reference-text"><cite id="CITEREFBanzhafGoodmanShenemanTrujillo2020" class="citation book cs1">Banzhaf, Wolfgang; Goodman, Erik; Sheneman, Leigh; Trujillo, Leonardo; Worzel, Bill, eds. (2020). <a rel="nofollow" class="external text" href="http://link.springer.com/10.1007/978-3-030-39958-0"><i>Genetic Programming Theory and Practice XVII</i></a>. Genetic and Evolutionary Computation. Cham: Springer International Publishing. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2F978-3-030-39958-0">10.1007/978-3-030-39958-0</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-3-030-39957-3</bdi>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:218531292">218531292</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230315193000/https://link.springer.com/book/10.1007/978-3-030-39958-0">Archived</a> from the original on March 15, 2023.</cite></span>
</li>
<li id="cite_note-128"><span class="mw-cite-backlink"><b><a href="#cite_ref-128">^</a></b></span> <span class="reference-text"><cite id="CITEREFWiblin2018" class="citation podcast cs1">Wiblin, Robert (October 2, 2018). <a rel="nofollow" class="external text" href="https://80000hours.org/podcast/episodes/paul-christiano-ai-alignment-solutions/">"Dr Paul Christiano on how OpenAI is developing real solutions to the 'AI alignment problem', and his vision of how humanity will progressively hand over decision-making to AI systems"</a> (Podcast). 80,000 hours. No.&nbsp;44. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221214050326/https://80000hours.org/podcast/episodes/paul-christiano-ai-alignment-solutions/">Archived</a> from the original on December 14, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-129"><span class="mw-cite-backlink"><b><a href="#cite_ref-129">^</a></b></span> <span class="reference-text"><cite id="CITEREFLehmanCluneMisevicAdami2020" class="citation journal cs1">Lehman, Joel; Clune, Jeff; Misevic, Dusan; Adami, Christoph; Altenberg, Lee; Beaulieu, Julie; Bentley, Peter J.; Bernard, Samuel; Beslon, Guillaume; Bryson, David M.; Cheney, Nick (2020). <a rel="nofollow" class="external text" href="https://direct.mit.edu/artl/article/26/2/274-306/93255">"The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities"</a>. <i>Artificial Life</i>. <b>26</b> (2): <span class="nowrap">274–</span>306. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1162%2Fartl_a_00319">10.1162/artl_a_00319</a></span>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://hdl.handle.net/10044%2F1%2F83343">10044/1/83343</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1064-5462">1064-5462</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/32271631">32271631</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:4519185">4519185</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20221010143108/https://direct.mit.edu/artl/article/26/2/274-306/93255">Archived</a> from the original on October 10, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-saavrm-130"><span class="mw-cite-backlink">^ <a href="#cite_ref-saavrm_130-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-saavrm_130-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLeikeKruegerEverittMartic2018" class="citation arxiv cs1">Leike, Jan; Krueger, David; Everitt, Tom; Martic, Miljan; Maini, Vishal; Legg, Shane (November 19, 2018). "Scalable agent alignment via reward modeling: a research direction". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1811.07871">1811.07871</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-OpenAIApproach-131"><span class="mw-cite-backlink">^ <a href="#cite_ref-OpenAIApproach_131-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-OpenAIApproach_131-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLeikeSchulmanWu2022" class="citation web cs1">Leike, Jan; Schulman, John; Wu, Jeffrey (August 24, 2022). <a rel="nofollow" class="external text" href="https://openai.com/blog/our-approach-to-alignment-research/">"Our approach to alignment research"</a>. <i>OpenAI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230215193559/https://openai.com/blog/our-approach-to-alignment-research/">Archived</a> from the original on February 15, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 9,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-132"><span class="mw-cite-backlink"><b><a href="#cite_ref-132">^</a></b></span> <span class="reference-text"><cite id="CITEREFWiggers2021" class="citation web cs1">Wiggers, Kyle (September 23, 2021). <a rel="nofollow" class="external text" href="https://venturebeat.com/2021/09/23/openai-unveils-model-that-can-summarize-books-of-any-length/">"OpenAI unveils model that can summarize books of any length"</a>. <i>VentureBeat</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220723215104/https://venturebeat.com/2021/09/23/openai-unveils-model-that-can-summarize-books-of-any-length/">Archived</a> from the original on July 23, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-133"><span class="mw-cite-backlink"><b><a href="#cite_ref-133">^</a></b></span> <span class="reference-text"><cite id="CITEREFSaundersYehWuBills2022" class="citation arxiv cs1">Saunders, William; Yeh, Catherine; Wu, Jeff; Bills, Steven; Ouyang, Long; Ward, Jonathan; Leike, Jan (June 13, 2022). "Self-critiquing models for assisting human evaluators". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2206.05802">2206.05802</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite>
<ul><li><cite id="CITEREFBaiKadavathKunduAskell2022" class="citation arxiv cs1">Bai, Yuntao; Kadavath, Saurav; Kundu, Sandipan; Askell, Amanda; Kernion, Jackson; Jones, Andy; Chen, Anna; Goldie, Anna; Mirhoseini, Azalia; McKinnon, Cameron; Chen, Carol; Olsson, Catherine; Olah, Christopher; Hernandez, Danny; Drain, Dawn (December 15, 2022). "Constitutional AI: Harmlessness from AI Feedback". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2212.08073">2212.08073</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></li></ul>
</span></li>
<li id="cite_note-134"><span class="mw-cite-backlink"><b><a href="#cite_ref-134">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://openai.com/blog/introducing-superalignment">"Introducing Superalignment"</a>. <i>openai.com</i><span class="reference-accessdate">. Retrieved <span class="nowrap">July 17,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-Falsehoods-135"><span class="mw-cite-backlink">^ <a href="#cite_ref-Falsehoods_135-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Falsehoods_135-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Falsehoods_135-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFWiggers2021" class="citation web cs1">Wiggers, Kyle (September 20, 2021). <a rel="nofollow" class="external text" href="https://venturebeat.com/2021/09/20/falsehoods-more-likely-with-large-language-models/">"Falsehoods more likely with large language models"</a>. <i>VentureBeat</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220804142703/https://venturebeat.com/2021/09/20/falsehoods-more-likely-with-large-language-models/">Archived</a> from the original on August 4, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-136"><span class="mw-cite-backlink"><b><a href="#cite_ref-136">^</a></b></span> <span class="reference-text"><cite id="CITEREFThe_Guardian2020" class="citation news cs1">The Guardian (September 8, 2020). <a rel="nofollow" class="external text" href="https://www.theguardian.com/commentisfree/2020/sep/08/robot-wrote-this-article-gpt-3">"A robot wrote this entire article. Are you scared yet, human?"</a>. <i>The Guardian</i>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0261-3077">0261-3077</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20200908090812/https://www.theguardian.com/commentisfree/2020/sep/08/robot-wrote-this-article-gpt-3">Archived</a> from the original on September 8, 2020<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite>
<ul><li><cite id="CITEREFHeaven2020" class="citation web cs1">Heaven, Will Douglas (July 20, 2020). <a rel="nofollow" class="external text" href="https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-gpt-3-nlp/">"OpenAI's new language generator GPT-3 is shockingly good—and completely mindless"</a>. <i>MIT Technology Review</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20200725175436/https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-gpt-3-nlp/">Archived</a> from the original on July 25, 2020<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-TruthfulAI-137"><span class="mw-cite-backlink">^ <a href="#cite_ref-TruthfulAI_137-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-TruthfulAI_137-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFEvansCotton-BarrattFinnvedenBales2021" class="citation arxiv cs1">Evans, Owain; Cotton-Barratt, Owen; Finnveden, Lukas; Bales, Adam; Balwit, Avital; Wills, Peter; Righetti, Luca; Saunders, William (October 13, 2021). "Truthful AI: Developing and governing AI that does not lie". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2110.06674">2110.06674</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-138"><span class="mw-cite-backlink"><b><a href="#cite_ref-138">^</a></b></span> <span class="reference-text"><cite id="CITEREFAlford2021" class="citation web cs1">Alford, Anthony (July 13, 2021). <a rel="nofollow" class="external text" href="https://www.infoq.com/news/2021/07/eleutherai-gpt-j/">"EleutherAI Open-Sources Six Billion Parameter GPT-3 Clone GPT-J"</a>. <i>InfoQ</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.infoq.com/news/2021/07/eleutherai-gpt-j/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite>
<ul><li><cite id="CITEREFRaeBorgeaudCaiMillican2022" class="citation arxiv cs1">Rae, Jack W.; Borgeaud, Sebastian; Cai, Trevor; Millican, Katie; Hoffmann, Jordan; Song, Francis; Aslanides, John; Henderson, Sarah; Ring, Roman; Young, Susannah; Rutherford, Eliza; Hennigan, Tom; Menick, Jacob; Cassirer, Albin; Powell, Richard (January 21, 2022). "Scaling Language Models: Methods, Analysis &amp; Insights from Training Gopher". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2112.11446">2112.11446</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></li></ul>
</span></li>
<li id="cite_note-139"><span class="mw-cite-backlink"><b><a href="#cite_ref-139">^</a></b></span> <span class="reference-text"><cite id="CITEREFNakanoHiltonBalajiWu2022" class="citation arxiv cs1">Nakano, Reiichiro; Hilton, Jacob; <a href="Suchir_Balaji" title="Suchir Balaji">Balaji, Suchir</a>; Wu, Jeff; Ouyang, Long; Kim, Christina; Hesse, Christopher; Jain, Shantanu; Kosaraju, Vineet; Saunders, William; Jiang, Xu; Cobbe, Karl; Eloundou, Tyna; Krueger, Gretchen; Button, Kevin (June 1, 2022). "WebGPT: Browser-assisted question-answering with human feedback". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2112.09332">2112.09332</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite>
<ul><li><cite id="CITEREFKumar2021" class="citation web cs1">Kumar, Nitish (December 23, 2021). <a rel="nofollow" class="external text" href="https://www.marktechpost.com/2021/12/22/openai-researchers-find-ways-to-more-accurately-answer-open-ended-questions-using-a-text-based-web-browser/">"OpenAI Researchers Find Ways To More Accurately Answer Open-Ended Questions Using A Text-Based Web Browser"</a>. <i>MarkTechPost</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.marktechpost.com/2021/12/22/openai-researchers-find-ways-to-more-accurately-answer-open-ended-questions-using-a-text-based-web-browser/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></li>
<li><cite id="CITEREFMenickTrebaczMikulikAslanides2022" class="citation journal cs1">Menick, Jacob; Trebacz, Maja; Mikulik, Vladimir; Aslanides, John; Song, Francis; Chadwick, Martin; Glaese, Mia; Young, Susannah; Campbell-Gillingham, Lucy; Irving, Geoffrey; McAleese, Nat (March 21, 2022). <a rel="nofollow" class="external text" href="https://www.deepmind.com/publications/gophercite-teaching-language-models-to-support-answers-with-verified-quotes">"Teaching language models to support answers with verified quotes"</a>. <i>DeepMind</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2203.11147">2203.11147</a></span>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.deepmind.com/publications/gophercite-teaching-language-models-to-support-answers-with-verified-quotes">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 12,</span> 2022</span>.</cite></li></ul>
</span></li>
<li id="cite_note-140"><span class="mw-cite-backlink"><b><a href="#cite_ref-140">^</a></b></span> <span class="reference-text"><cite id="CITEREFAskellBaiChenDrain2021" class="citation arxiv cs1">Askell, Amanda; Bai, Yuntao; Chen, Anna; Drain, Dawn; Ganguli, Deep; Henighan, Tom; Jones, Andy; Joseph, Nicholas; Mann, Ben; DasSarma, Nova; Elhage, Nelson; Hatfield-Dodds, Zac; Hernandez, Danny; Kernion, Jackson; Ndousse, Kamal (December 9, 2021). "A General Language Assistant as a Laboratory for Alignment". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2112.00861">2112.00861</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-141"><span class="mw-cite-backlink"><b><a href="#cite_ref-141">^</a></b></span> <span class="reference-text"><cite id="CITEREFCox2023" class="citation web cs1">Cox, Joseph (March 15, 2023). <a rel="nofollow" class="external text" href="https://www.vice.com/en/article/gpt4-hired-unwitting-taskrabbit-worker/">"GPT-4 Hired Unwitting TaskRabbit Worker By Pretending to Be 'Vision-Impaired' Human"</a>. <i>Vice</i><span class="reference-accessdate">. Retrieved <span class="nowrap">April 10,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-142"><span class="mw-cite-backlink"><b><a href="#cite_ref-142">^</a></b></span> <span class="reference-text"><cite id="CITEREFScheurerBalesniHobbhahn2023" class="citation arxiv cs1">Scheurer, Jérémy; Balesni, Mikita; Hobbhahn, Marius (2023). "Technical Report: Large Language Models can Strategically Deceive their Users when Put Under Pressure". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2311.07590">2311.07590</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite></span>
</li>
<li id="cite_note-143"><span class="mw-cite-backlink"><b><a href="#cite_ref-143">^</a></b></span> <span class="reference-text"><cite id="CITEREFKentonEverittWeidingerGabriel2021" class="citation web cs1">Kenton, Zachary; Everitt, Tom; Weidinger, Laura; Gabriel, Iason; Mikulik, Vladimir; Irving, Geoffrey (March 30, 2021). <a rel="nofollow" class="external text" href="https://deepmindsafetyresearch.medium.com/alignment-of-language-agents-9fbc7dd52c6c">"Alignment of Language Agents"</a>. <i>DeepMind Safety Research – Medium</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114142/https://deepmindsafetyresearch.medium.com/alignment-of-language-agents-9fbc7dd52c6c">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">July 23,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-144"><span class="mw-cite-backlink"><b><a href="#cite_ref-144">^</a></b></span> <span class="reference-text"><cite id="CITEREFParkGoldsteinO’GaraChen2024" class="citation journal cs1">Park, Peter S.; Goldstein, Simon; O’Gara, Aidan; Chen, Michael; Hendrycks, Dan (May 2024). <a rel="nofollow" class="external text" href="https://doi.org/10.1016/j.patter.2024.100988">"AI deception: A survey of examples, risks, and potential solutions"</a>. <i>Patterns</i>. <b>5</b> (5): 100988. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.patter.2024.100988">10.1016/j.patter.2024.100988</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2666-3899">2666-3899</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a>&nbsp;<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11117051">11117051</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/38800366">38800366</a>.</cite></span>
</li>
<li id="cite_note-145"><span class="mw-cite-backlink"><b><a href="#cite_ref-145">^</a></b></span> <span class="reference-text"><cite id="CITEREFZia2025" class="citation web cs1">Zia, Tehseen (January 7, 2025). <a rel="nofollow" class="external text" href="https://www.unite.ai/can-ai-be-trusted-the-challenge-of-alignment-faking/">"Can AI Be Trusted? The Challenge of Alignment Faking"</a>. <i>Unite.AI</i><span class="reference-accessdate">. Retrieved <span class="nowrap">February 18,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-:6-146"><span class="mw-cite-backlink">^ <a href="#cite_ref-:6_146-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:6_146-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFPerrigo2024" class="citation magazine cs1">Perrigo, Billy (December 18, 2024). <a rel="nofollow" class="external text" href="https://time.com/7202784/ai-research-strategic-lying/">"Exclusive: New Research Shows AI Strategically Lying"</a>. <i>TIME</i><span class="reference-accessdate">. Retrieved <span class="nowrap">February 18,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-147"><span class="mw-cite-backlink"><b><a href="#cite_ref-147">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.anthropic.com/research/alignment-faking">"Alignment faking in large language models"</a>. <i>Anthropic</i>. December 18, 2024<span class="reference-accessdate">. Retrieved <span class="nowrap">February 17,</span> 2025</span>.</cite></span>
</li>
<li id="cite_note-148"><span class="mw-cite-backlink"><b><a href="#cite_ref-148">^</a></b></span> <span class="reference-text"><cite id="CITEREFGreenblattDenisonWrightRoger2024" class="citation arxiv cs1">Greenblatt, Ryan; Denison, Carson; Wright, Benjamin; Roger, Fabien; MacDiarmid, Monte; Marks, Sam; Treutlein, Johannes; Belonax, Tim; Chen, Jack (December 20, 2024). "Alignment faking in large language models". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2412.14093">2412.14093</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-149"><span class="mw-cite-backlink"><b><a href="#cite_ref-149">^</a></b></span> <span class="reference-text"><cite id="CITEREFMcCarthyMinskyRochesterShannon2006" class="citation journal cs1">McCarthy, John; Minsky, Marvin L.; Rochester, Nathaniel; Shannon, Claude E. (December 15, 2006). <a rel="nofollow" class="external text" href="https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/1904">"A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955"</a>. <i>AI Magazine</i>. <b>27</b> (4): 12. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1609%2Faimag.v27i4.1904">10.1609/aimag.v27i4.1904</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2371-9621">2371-9621</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:19439915">19439915</a>.</cite></span>
</li>
<li id="cite_note-150"><span class="mw-cite-backlink"><b><a href="#cite_ref-150">^</a></b></span> <span class="reference-text"><cite id="CITEREFWangMaFengZhang2024" class="citation cs2">Wang, Lei; Ma, Chen; Feng, Xueyang; Zhang, Zeyu; Yang, Hao; Zhang, Jingsen; Chen, Zhiyuan; Tang, Jiakai; Chen, Xu (2024), "A survey on large language model based autonomous agents", <i>Frontiers of Computer Science</i>, <b>18</b> (6) 186345, <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2308.11432">2308.11432</a></span>, <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11704-024-40231-1">10.1007/s11704-024-40231-1</a></cite></span>
</li>
<li id="cite_note-151"><span class="mw-cite-backlink"><b><a href="#cite_ref-151">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://fortune.com/2023/05/02/godfather-ai-geoff-hinton-google-warns-artificial-intelligence-nightmare-scenario/">"'The Godfather of A.I.' warns of 'nightmare scenario' where artificial intelligence begins to seek power"</a>. <i>Fortune</i><span class="reference-accessdate">. Retrieved <span class="nowrap">May 4,</span> 2023</span>.</cite>
<ul><li><cite class="citation news cs1"><a rel="nofollow" class="external text" href="https://www.technologyreview.com/2016/11/02/156285/yes-we-are-worried-about-the-existential-risk-of-artificial-intelligence/">"Yes, We Are Worried About the Existential Risk of Artificial Intelligence"</a>. <i>MIT Technology Review</i><span class="reference-accessdate">. Retrieved <span class="nowrap">May 4,</span> 2023</span>.</cite></li></ul>
</span></li>
<li id="cite_note-quanta-hide-seek2-152"><span class="mw-cite-backlink"><b><a href="#cite_ref-quanta-hide-seek2_152-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFOrnes2019" class="citation web cs1">Ornes, Stephen (November 18, 2019). <a rel="nofollow" class="external text" href="https://www.quantamagazine.org/artificial-intelligence-discovers-tool-use-in-hide-and-seek-games-20191118/">"Playing Hide-and-Seek, Machines Invent New Tools"</a>. <i>Quanta Magazine</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.quantamagazine.org/artificial-intelligence-discovers-tool-use-in-hide-and-seek-games-20191118/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-153"><span class="mw-cite-backlink"><b><a href="#cite_ref-153">^</a></b></span> <span class="reference-text"><cite id="CITEREFBakerKanitscheiderMarkovWu2019" class="citation web cs1">Baker, Bowen; Kanitscheider, Ingmar; Markov, Todor; Wu, Yi; Powell, Glenn; McGrew, Bob; Mordatch, Igor (September 17, 2019). <a rel="nofollow" class="external text" href="https://openai.com/blog/emergent-tool-use/">"Emergent Tool Use from Multi-Agent Interaction"</a>. <i>OpenAI</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20220925043450/https://openai.com/blog/emergent-tool-use/">Archived</a> from the original on September 25, 2022<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-154"><span class="mw-cite-backlink"><b><a href="#cite_ref-154">^</a></b></span> <span class="reference-text"><cite id="CITEREFLuLuLangeFoerster2024" class="citation arxiv cs1">Lu, Chris; Lu, Cong; Lange, Robert Tjarko; Foerster, Jakob; Clune, Jeff; Ha, David (August 15, 2024). "The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2408.06292">2408.06292</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>]. <q>In some cases, when The AI Scientist's experiments exceeded our imposed time limits, it attempted to edit the code to extend the time limit arbitrarily</q></cite></span>
</li>
<li id="cite_note-155"><span class="mw-cite-backlink"><b><a href="#cite_ref-155">^</a></b></span> <span class="reference-text"><cite id="CITEREFEdwards2024" class="citation web cs1">Edwards, Benj (August 14, 2024). <a rel="nofollow" class="external text" href="https://arstechnica.com/information-technology/2024/08/research-ai-model-unexpectedly-modified-its-own-code-to-extend-runtime/">"Research AI model unexpectedly modified its own code to extend runtime"</a>. <i>Ars Technica</i><span class="reference-accessdate">. Retrieved <span class="nowrap">August 19,</span> 2024</span>.</cite></span>
</li>
<li id="cite_note-156"><span class="mw-cite-backlink"><b><a href="#cite_ref-156">^</a></b></span> <span class="reference-text"><cite id="CITEREFShermer2017" class="citation web cs1">Shermer, Michael (March 1, 2017). <a rel="nofollow" class="external text" href="https://www.scientificamerican.com/article/artificial-intelligence-is-not-a-threat-mdash-yet/">"Artificial Intelligence Is Not a Threat—Yet"</a>. <i>Scientific American</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20171201051401/https://www.scientificamerican.com/article/artificial-intelligence-is-not-a-threat-mdash-yet/">Archived</a> from the original on December 1, 2017<span class="reference-accessdate">. Retrieved <span class="nowrap">August 26,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-158"><span class="mw-cite-backlink"><b><a href="#cite_ref-158">^</a></b></span> <span class="reference-text"><cite id="CITEREFBrownMannRyderSubbiah2020" class="citation arxiv cs1">Brown, Tom B.; Mann, Benjamin; Ryder, Nick; Subbiah, Melanie; Kaplan, Jared; Dhariwal, Prafulla; Neelakantan, Arvind; Shyam, Pranav; Sastry, Girish; Askell, Amanda; Agarwal, Sandhini; Herbert-Voss, Ariel; Krueger, Gretchen; Henighan, Tom; Child, Rewon (July 22, 2020). "Language Models are Few-Shot Learners". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2005.14165">2005.14165</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CL">cs.CL</a>].</cite>
<ul><li><cite id="CITEREFLaskinWangOhParisotto2022" class="citation arxiv cs1">Laskin, Michael; Wang, Luyu; Oh, Junhyuk; Parisotto, Emilio; Spencer, Stephen; Steigerwald, Richie; Strouse, D. J.; Hansen, Steven; Filos, Angelos; Brooks, Ethan; Gazeau, Maxime; Sahni, Himanshu; Singh, Satinder; Mnih, Volodymyr (October 25, 2022). "In-context Reinforcement Learning with Algorithm Distillation". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.14215">2210.14215</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></li></ul>
</span></li>
<li id="cite_note-159"><span class="mw-cite-backlink"><b><a href="#cite_ref-159">^</a></b></span> <span class="reference-text"><cite id="CITEREFMeloMáximoSomaCastro2025" class="citation journal cs1">Melo, Gabriel A.; Máximo, Marcos R. O. A.; Soma, Nei Y.; Castro, Paulo A. L. (2025). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12050267">"Machines that halt resolve the undecidability of artificial intelligence alignment"</a>. <i>Scientific Reports</i>. <b>15</b> (1): 15591. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2025NatSR..1515591M">2025NatSR..1515591M</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1038%2Fs41598-025-99060-2">10.1038/s41598-025-99060-2</a>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a>&nbsp;<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12050267">12050267</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/40320467">40320467</a>.</cite></span>
</li>
<li id="cite_note-GoalMisgeneralization-160"><span class="mw-cite-backlink">^ <a href="#cite_ref-GoalMisgeneralization_160-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-GoalMisgeneralization_160-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-GoalMisgeneralization_160-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFShahVarmaKumarPhuong2022" class="citation journal cs1">Shah, Rohin; Varma, Vikrant; Kumar, Ramana; Phuong, Mary; Krakovna, Victoria; Uesato, Jonathan; Kenton, Zac (November 2, 2022). <a rel="nofollow" class="external text" href="https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924">"Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals"</a>. <i>Medium</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.01790">2210.01790</a></span><span class="reference-accessdate">. Retrieved <span class="nowrap">April 2,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-rloamls-161"><span class="mw-cite-backlink">^ <a href="#cite_ref-rloamls_161-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-rloamls_161-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFHubingervan_MerwijkMikulikSkalse2021" class="citation arxiv cs1">Hubinger, Evan; van Merwijk, Chris; Mikulik, Vladimir; Skalse, Joar; Garrabrant, Scott (December 1, 2021). "Risks from Learned Optimization in Advanced Machine Learning Systems". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1906.01820">1906.01820</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-162"><span class="mw-cite-backlink"><b><a href="#cite_ref-162">^</a></b></span> <span class="reference-text"><cite id="CITEREFZhangChanYanBose2022" class="citation journal cs1">Zhang, Xiaoge; Chan, Felix T.S.; Yan, Chao; Bose, Indranil (2022). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://linkinghub.elsevier.com/retrieve/pii/S0167923622000719">"Towards risk-aware artificial intelligence and machine learning systems: An overview"</a></span>. <i>Decision Support Systems</i>. <b>159</b> 113800. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.dss.2022.113800">10.1016/j.dss.2022.113800</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:248585546">248585546</a>.</cite></span>
</li>
<li id="cite_note-163"><span class="mw-cite-backlink"><b><a href="#cite_ref-163">^</a></b></span> <span class="reference-text"><cite id="CITEREFDemskiGarrabrant2020" class="citation arxiv cs1">Demski, Abram; Garrabrant, Scott (October 6, 2020). "Embedded Agency". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1902.09469">1902.09469</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-causal_influence2-164"><span class="mw-cite-backlink">^ <a href="#cite_ref-causal_influence2_164-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-causal_influence2_164-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFEverittOrtegaBarnesLegg2019" class="citation arxiv cs1">Everitt, Tom; Ortega, Pedro A.; Barnes, Elizabeth; Legg, Shane (September 6, 2019). "Understanding Agent Incentives using Causal Influence Diagrams. Part I: Single Action Settings". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1902.09980">1902.09980</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-:323-165"><span class="mw-cite-backlink">^ <a href="#cite_ref-:323_165-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:323_165-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFCohenHutterOsborne2022" class="citation journal cs1">Cohen, Michael K.; Hutter, Marcus; Osborne, Michael A. (August 29, 2022). <a rel="nofollow" class="external text" href="https://onlinelibrary.wiley.com/doi/10.1002/aaai.12064">"Advanced artificial agents intervene in the provision of reward"</a>. <i>AI Magazine</i>. <b>43</b> (3): <span class="nowrap">282–</span>293. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1002%2Faaai.12064">10.1002/aaai.12064</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0738-4602">0738-4602</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:235489158">235489158</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210153534/https://onlinelibrary.wiley.com/doi/10.1002/aaai.12064">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">September 6,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-Hadfield-Menell2019-166"><span class="mw-cite-backlink"><b><a href="#cite_ref-Hadfield-Menell2019_166-0">^</a></b></span> <span class="reference-text"> <cite id="CITEREFHadfield-MenellHadfield2019" class="citation conference cs1">Hadfield-Menell, Dylan; Hadfield, Gillian K (2019). "Incomplete contracting and AI alignment". <i>Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society</i>. pp.&nbsp;<span class="nowrap">417–</span>422.</cite></span>
</li>
<li id="cite_note-Hanson2019-167"><span class="mw-cite-backlink"><b><a href="#cite_ref-Hanson2019_167-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFHanson2019" class="citation web cs1">Hanson, Robin (April 10, 2019). <a rel="nofollow" class="external text" href="https://www.overcomingbias.com/p/agency-failure-ai-apocalypsehtml">"Agency Failure or AI Apocalypse?"</a>. <i>Overcoming Bias</i><span class="reference-accessdate">. Retrieved <span class="nowrap">September 20,</span> 2023</span>.</cite></span>
</li>
<li id="cite_note-168"><span class="mw-cite-backlink"><b><a href="#cite_ref-168">^</a></b></span> <span class="reference-text"><cite id="CITEREFHamilton2020" class="citation cs2">Hamilton, Andy (2020), <a rel="nofollow" class="external text" href="https://plato.stanford.edu/entries/conservatism/">"Conservatism"</a>, in Zalta, Edward N. (ed.), <i>The Stanford Encyclopedia of Philosophy</i> (Spring 2020&nbsp;ed.), Metaphysics Research Lab, Stanford University<span class="reference-accessdate">, retrieved <span class="nowrap">October 16,</span> 2024</span></cite></span>
</li>
<li id="cite_note-169"><span class="mw-cite-backlink"><b><a href="#cite_ref-169">^</a></b></span> <span class="reference-text"><cite id="CITEREFTaylorYudkowskyLaVictoireCritch2016" class="citation web cs1">Taylor, Jessica; Yudkowsky, Eliezer; LaVictoire, Patrick; Critch, Andrew (July 27, 2016). <a rel="nofollow" class="external text" href="https://intelligence.org/files/AlignmentMachineLearning.pdf">"Alignment for Advanced Machine Learning Systems"</a> <span class="cs1-format">(PDF)</span>.</cite></span>
</li>
<li id="cite_note-170"><span class="mw-cite-backlink"><b><a href="#cite_ref-170">^</a></b></span> <span class="reference-text"><cite id="CITEREFBengio2024" class="citation web cs1">Bengio, Yoshua (February 26, 2024). <a rel="nofollow" class="external text" href="https://yoshuabengio.org/2024/02/26/towards-a-cautious-scientist-ai-with-convergent-safety-bounds/">"Towards a Cautious Scientist AI with Convergent Safety Bounds"</a>.</cite></span>
</li>
<li id="cite_note-171"><span class="mw-cite-backlink"><b><a href="#cite_ref-171">^</a></b></span> <span class="reference-text"><cite id="CITEREFCohenHutter2020" class="citation journal cs1">Cohen, Michael; Hutter, Marcus (2020). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v125/cohen20a/cohen20a.pdf">"Pessimism about unknown unknowns inspires conservatism"</a> <span class="cs1-format">(PDF)</span>. <i>Proceedings of Machine Learning Research</i>. <b>125</b>: <span class="nowrap">1344–</span>1373. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2006.08753">2006.08753</a></span>.</cite></span>
</li>
<li id="cite_note-172"><span class="mw-cite-backlink"><b><a href="#cite_ref-172">^</a></b></span> <span class="reference-text"><cite id="CITEREFLiuReyzinZiebart2015" class="citation journal cs1">Liu, Anqi; Reyzin, Lev; Ziebart, Brian (February 21, 2015). <a rel="nofollow" class="external text" href="https://ojs.aaai.org/index.php/AAAI/article/view/9609">"Shift-Pessimistic Active Learning Using Robust Bias-Aware Prediction"</a>. <i>Proceedings of the AAAI Conference on Artificial Intelligence</i>. <b>29</b> (1). <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1609%2Faaai.v29i1.9609">10.1609/aaai.v29i1.9609</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2374-3468">2374-3468</a>.</cite></span>
</li>
<li id="cite_note-173"><span class="mw-cite-backlink"><b><a href="#cite_ref-173">^</a></b></span> <span class="reference-text"><cite id="CITEREFLiuShenCuiZhou2021" class="citation journal cs1">Liu, Jiashuo; Shen, Zheyan; Cui, Peng; Zhou, Linjun; Kuang, Kun; Li, Bo; Lin, Yishi (May 18, 2021). <a rel="nofollow" class="external text" href="https://ojs.aaai.org/index.php/AAAI/article/view/17050">"Stable Adversarial Learning under Distributional Shifts"</a>. <i>Proceedings of the AAAI Conference on Artificial Intelligence</i>. <b>35</b> (10): <span class="nowrap">8662–</span>8670. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2006.04414">2006.04414</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1609%2Faaai.v35i10.17050">10.1609/aaai.v35i10.17050</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2374-3468">2374-3468</a>.</cite></span>
</li>
<li id="cite_note-174"><span class="mw-cite-backlink"><b><a href="#cite_ref-174">^</a></b></span> <span class="reference-text"><cite id="CITEREFRoyXuPokutta2017" class="citation journal cs1">Roy, Aurko; Xu, Huan; Pokutta, Sebastian (2017). <a rel="nofollow" class="external text" href="https://papers.nips.cc/paper_files/paper/2017/hash/84c6494d30851c63a55cdb8cb047fadd-Abstract.html">"Reinforcement Learning under Model Mismatch"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>30</b>. Curran Associates, Inc.</cite></span>
</li>
<li id="cite_note-175"><span class="mw-cite-backlink"><b><a href="#cite_ref-175">^</a></b></span> <span class="reference-text"><cite id="CITEREFPintoDavidsonSukthankarGupta2017" class="citation journal cs1">Pinto, Lerrel; Davidson, James; Sukthankar, Rahul; Gupta, Abhinav (July 17, 2017). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v70/pinto17a.html">"Robust Adversarial Reinforcement Learning"</a>. <i>Proceedings of the 34th International Conference on Machine Learning</i>. PMLR: <span class="nowrap">2817–</span>2826.</cite></span>
</li>
<li id="cite_note-176"><span class="mw-cite-backlink"><b><a href="#cite_ref-176">^</a></b></span> <span class="reference-text"><cite id="CITEREFWangZou2021" class="citation journal cs1">Wang, Yue; Zou, Shaofeng (2021). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper/2021/hash/3a4496776767aaa99f9804d0905fe584-Abstract.html">"Online Robust Reinforcement Learning with Model Uncertainty"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>34</b>. Curran Associates, Inc.: <span class="nowrap">7193–</span>7206. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2109.14523">2109.14523</a></span>.</cite></span>
</li>
<li id="cite_note-177"><span class="mw-cite-backlink"><b><a href="#cite_ref-177">^</a></b></span> <span class="reference-text"><cite id="CITEREFBlanchetLuZhangZhong2023" class="citation journal cs1">Blanchet, Jose; Lu, Miao; Zhang, Tong; Zhong, Han (December 15, 2023). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2023/hash/d31b005d817e9c635ec8ffb0fb90190e-Abstract-Conference.html">"Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial Coverage"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>36</b>: <span class="nowrap">66845–</span>66859. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2305.09659">2305.09659</a></span>.</cite></span>
</li>
<li id="cite_note-178"><span class="mw-cite-backlink"><b><a href="#cite_ref-178">^</a></b></span> <span class="reference-text"><cite id="CITEREFLevineKumarTuckerFu2020" class="citation arxiv cs1">Levine, Sergey; Kumar, Aviral; Tucker, George; Fu, Justin (November 1, 2020). "Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2005.01643">2005.01643</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-179"><span class="mw-cite-backlink"><b><a href="#cite_ref-179">^</a></b></span> <span class="reference-text"><cite id="CITEREFRigterLacerdaHawes2022" class="citation journal cs1">Rigter, Marc; Lacerda, Bruno; Hawes, Nick (December 6, 2022). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2022/hash/6691c5e4a199b72dffd9c90acb63bcd6-Abstract-Conference.html">"RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>35</b>: <span class="nowrap">16082–</span>16097. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2204.12581">2204.12581</a></span>.</cite></span>
</li>
<li id="cite_note-180"><span class="mw-cite-backlink"><b><a href="#cite_ref-180">^</a></b></span> <span class="reference-text"><cite id="CITEREFGuoYunfengGeng2022" class="citation journal cs1">Guo, Kaiyang; Yunfeng, Shao; Geng, Yanhui (December 6, 2022). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2022/hash/03469b1a66e351b18272be23baf3b809-Abstract-Conference.html">"Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics Belief"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>35</b>: <span class="nowrap">449–</span>461. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2210.06692">2210.06692</a></span>.</cite></span>
</li>
<li id="cite_note-181"><span class="mw-cite-backlink"><b><a href="#cite_ref-181">^</a></b></span> <span class="reference-text"><cite id="CITEREFCosteAnwarKirkKrueger2024" class="citation journal cs1">Coste, Thomas; Anwar, Usman; Kirk, Robert; Krueger, David (January 16, 2024). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=dcjtMYkpXx">"Reward Model Ensembles Help Mitigate Overoptimization"</a>. <i>International Conference on Learning Representations</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2310.02743">2310.02743</a></span>.</cite></span>
</li>
<li id="cite_note-182"><span class="mw-cite-backlink"><b><a href="#cite_ref-182">^</a></b></span> <span class="reference-text"><cite id="CITEREFLiuLuZhangLiu2024" class="citation arxiv cs1">Liu, Zhihan; Lu, Miao; Zhang, Shenao; Liu, Boyi; Guo, Hongyi; Yang, Yingxiang; Blanchet, Jose; Wang, Zhaoran (May 26, 2024). "Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2405.16436">2405.16436</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-183"><span class="mw-cite-backlink"><b><a href="#cite_ref-183">^</a></b></span> <span class="reference-text"><cite id="CITEREFCohenHutterNanda2022" class="citation journal cs1">Cohen, Michael K.; Hutter, Marcus; Nanda, Neel (2022). <a rel="nofollow" class="external text" href="https://jmlr.org/papers/v23/21-0618.html">"Fully General Online Imitation Learning"</a>. <i>Journal of Machine Learning Research</i>. <b>23</b> (334): <span class="nowrap">1–</span>30. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2102.08686">2102.08686</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1533-7928">1533-7928</a>.</cite></span>
</li>
<li id="cite_note-184"><span class="mw-cite-backlink"><b><a href="#cite_ref-184">^</a></b></span> <span class="reference-text"><cite id="CITEREFChangUeharaSreenivasKidambi2021" class="citation journal cs1">Chang, Jonathan; Uehara, Masatoshi; Sreenivas, Dhruv; Kidambi, Rahul; Sun, Wen (2021). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2021/hash/07d5938693cc3903b261e1a3844590ed-Abstract.html">"Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial Coverage"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>34</b>. Curran Associates, Inc.: <span class="nowrap">965–</span>979.</cite></span>
</li>
<li id="cite_note-185"><span class="mw-cite-backlink"><b><a href="#cite_ref-185">^</a></b></span> <span class="reference-text"><cite id="CITEREFBoydVandenberghe2023" class="citation book cs1">Boyd, Stephen P.; Vandenberghe, Lieven (2023). <i>Convex optimization</i> (Version 29&nbsp;ed.). Cambridge New York Melbourne New Delhi Singapore: Cambridge University Press. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0-521-83378-3</bdi>.</cite></span>
</li>
<li id="cite_note-186"><span class="mw-cite-backlink"><b><a href="#cite_ref-186">^</a></b></span> <span class="reference-text"><cite id="CITEREFKosoyAppel2021" class="citation journal cs1">Kosoy, Vanessa; Appel, Alexander (November 30, 2021). <a rel="nofollow" class="external text" href="https://www.alignmentforum.org/posts/gHgs2e2J5azvGFatb/infra-bayesian-physicalism-a-formal-theory-of-naturalized">"Infra-Bayesian physicalism: a formal theory of naturalized induction"</a>. <i>Alignment Forum</i>.</cite></span>
</li>
<li id="cite_note-187"><span class="mw-cite-backlink"><b><a href="#cite_ref-187">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://web.archive.org/web/20230216065407/https://www.un.org/en/content/common-agenda-report/">"UN Secretary-General's report on "Our Common Agenda""</a>. 2021. p.&nbsp;63. Archived from <a rel="nofollow" class="external text" href="https://www.un.org/en/content/common-agenda-report/">the original</a> on February 16, 2023. <q>[T]he Compact could also promote regulation of artificial intelligence to ensure that this is aligned with shared global values</q></cite></span>
</li>
<li id="cite_note-188"><span class="mw-cite-backlink"><b><a href="#cite_ref-188">^</a></b></span> <span class="reference-text"><cite id="CITEREFThe_National_New_Generation_Artificial_Intelligence_Governance_Specialist_Committee2021" class="citation web cs1">The National New Generation Artificial Intelligence Governance Specialist Committee (October 12, 2021) [2021-09-25]. <a rel="nofollow" class="external text" href="https://cset.georgetown.edu/publication/ethical-norms-for-new-generation-artificial-intelligence-released/">"Ethical Norms for New Generation Artificial Intelligence Released"</a>. Translated by <a href="Center_for_Security_and_Emerging_Technology" title="Center for Security and Emerging Technology">Center for Security and Emerging Technology</a>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114220/https://cset.georgetown.edu/publication/ethical-norms-for-new-generation-artificial-intelligence-released/">Archived</a> from the original on February 10, 2023.</cite></span>
</li>
<li id="cite_note-189"><span class="mw-cite-backlink"><b><a href="#cite_ref-189">^</a></b></span> <span class="reference-text"><cite id="CITEREFRichardson2021" class="citation news cs1">Richardson, Tim (September 22, 2021). <a rel="nofollow" class="external text" href="https://www.theregister.com/2021/09/22/uk_10_year_national_ai_strategy/">"UK publishes National Artificial Intelligence Strategy"</a>. <i>The Register</i>. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114137/https://www.theregister.com/2021/09/22/uk_10_year_national_ai_strategy/">Archived</a> from the original on February 10, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">November 14,</span> 2021</span>.</cite></span>
</li>
<li id="cite_note-190"><span class="mw-cite-backlink"><b><a href="#cite_ref-190">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114139/https://www.gov.uk/government/publications/national-ai-strategy/national-ai-strategy-html-version">"The National AI Strategy of the UK"</a>. 2021. Archived from <a rel="nofollow" class="external text" href="https://www.gov.uk/government/publications/national-ai-strategy/national-ai-strategy-html-version">the original</a> on February 10, 2023. <q>The government takes the long term risk of non-aligned Artificial General Intelligence, and the unforeseeable changes that it would mean for the UK and the world, seriously.</q></cite></span>
</li>
<li id="cite_note-191"><span class="mw-cite-backlink"><b><a href="#cite_ref-191">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://web.archive.org/web/20230210114139/https://www.gov.uk/government/publications/national-ai-strategy/national-ai-strategy-html-version">"The National AI Strategy of the UK"</a>. 2021. actions 9 and 10 of the section "Pillar 3 – Governing AI Effectively". Archived from <a rel="nofollow" class="external text" href="https://www.gov.uk/government/publications/national-ai-strategy/national-ai-strategy-html-version">the original</a> on February 10, 2023.</cite></span>
</li>
<li id="cite_note-192"><span class="mw-cite-backlink"><b><a href="#cite_ref-192">^</a></b></span> <span class="reference-text"><cite class="citation book cs1"><a rel="nofollow" class="external text" href="https://www.nscai.gov/wp-content/uploads/2021/03/Full-Report-Digital-1.pdf"><i>NSCAI Final Report</i></a> <span class="cs1-format">(PDF)</span>. Washington, DC: The National Security Commission on Artificial Intelligence. 2021. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230215110858/https://www.nscai.gov/wp-content/uploads/2021/03/Full-Report-Digital-1.pdf">Archived</a> <span class="cs1-format">(PDF)</span> from the original on February 15, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">October 17,</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-193"><span class="mw-cite-backlink"><b><a href="#cite_ref-193">^</a></b></span> <span class="reference-text"><cite id="CITEREFRobert_Lee_Poe2023" class="citation arxiv cs1">Robert Lee Poe (2023). "Why Fair Automated Hiring Systems Breach EU Non-Discrimination Law". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2311.03900">2311.03900</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CY">cs.CY</a>].</cite></span>
</li>
<li id="cite_note-194"><span class="mw-cite-backlink"><b><a href="#cite_ref-194">^</a></b></span> <span class="reference-text"><cite id="CITEREFDe_Vos2020" class="citation journal cs1">De Vos, Marc (2020). <a rel="nofollow" class="external text" href="https://doi.org/10.1177/1358229120927947">"The European Court of Justice and the march towards substantive equality in European Union anti-discrimination law"</a>. <i>International Journal of Discrimination and the Law</i>. <b>20</b>: <span class="nowrap">62–</span>87. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1177%2F1358229120927947">10.1177/1358229120927947</a>.</cite></span>
</li>
<li id="cite_note-195"><span class="mw-cite-backlink"><b><a href="#cite_ref-195">^</a></b></span> <span class="reference-text"><cite id="CITEREFIrvingAskell2016" class="citation journal cs1">Irving, Geoffrey; Askell, Amanda (June 9, 2016). "Chern number in Ising models with spatially modulated real and complex fields". <i>Physical Review A</i>. <b>94</b> (5): 052113. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1606.03535">1606.03535</a></span>. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2016PhRvA..94e2113L">2016PhRvA..94e2113L</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1103%2FPhysRevA.94.052113">10.1103/PhysRevA.94.052113</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:118699363">118699363</a>.</cite></span>
</li>
<li id="cite_note-196"><span class="mw-cite-backlink"><b><a href="#cite_ref-196">^</a></b></span> <span class="reference-text"><cite id="CITEREFMitelutSmithVamplew2023" class="citation arxiv cs1">Mitelut, Catalin; Smith, Ben; Vamplew, Peter (May 30, 2023). "Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2305.19223">2305.19223</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></span>
</li>
<li id="cite_note-197"><span class="mw-cite-backlink"><b><a href="#cite_ref-197">^</a></b></span> <span class="reference-text"><cite id="CITEREFGabriel2020" class="citation journal cs1">Gabriel, Iason (September 1, 2020). <a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11023-020-09539-2">"Artificial Intelligence, Values, and Alignment"</a>. <i>Minds and Machines</i>. <b>30</b> (3): <span class="nowrap">411–</span>437. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2001.09768">2001.09768</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1007%2Fs11023-020-09539-2">10.1007/s11023-020-09539-2</a></span>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:210920551">210920551</a>.</cite></span>
</li>
<li id="cite_note-198"><span class="mw-cite-backlink"><b><a href="#cite_ref-198">^</a></b></span> <span class="reference-text"><cite id="CITEREFRussell2019" class="citation book cs1">Russell, Stuart J. (2019). <a rel="nofollow" class="external text" href="https://www.penguinrandomhouse.com/books/566677/human-compatible-by-stuart-russell/"><i>Human Compatible: Artificial Intelligence and the Problem of Control</i></a>. Penguin Random House.</cite></span>
</li>
<li id="cite_note-199"><span class="mw-cite-backlink"><b><a href="#cite_ref-199">^</a></b></span> <span class="reference-text"><cite id="CITEREFDafoe2019" class="citation journal cs1">Dafoe, Allan (2019). <a rel="nofollow" class="external text" href="https://www.nature.com/articles/s41586-019-1420-6">"AI policy: A roadmap"</a>. <i>Nature</i>.</cite></span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="Further_reading">Further reading</h2></div>
<ul><li><cite id="CITEREFBrockman2019" class="citation book cs1"><a href="John_Brockman_(literary_agent)" title="John Brockman (literary agent)">Brockman, John</a>, ed. (2019). <a href="Possible_Minds" title="Possible Minds"><i>Possible Minds: Twenty-five Ways of Looking at AI</i></a> (Kindle&nbsp;ed.). Penguin Press. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-0525557999</bdi>.</cite><span class="cs1-maint citation-comment"><code class="cs1-code">{{cite book}}</code>: CS1 maint: ref duplicates default (link)</span></li>
<li><cite id="CITEREFNgoChanMindermann2023" class="citation arxiv cs1">Ngo, Richard; et&nbsp;al. (2023). "The Alignment Problem from a Deep Learning Perspective". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2209.00626">2209.00626</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></li>
<li><cite id="CITEREFJiQiuChen2023" class="citation arxiv cs1">Ji, Jiaming; et&nbsp;al. (2023). "AI Alignment: A Comprehensive Survey". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2310.19852">2310.19852</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.AI">cs.AI</a>].</cite></li></ul>
<div class="mw-heading mw-heading2"><h2 id="External_links">External links</h2></div>
<ul><li><a rel="nofollow" class="external text" href="https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml">Specification gaming examples in AI</a>, via <a rel="nofollow" class="external text" href="https://deepmind.com/blog/article/Specification-gaming-the-flip-side-of-AI-ingenuity">DeepMind</a></li></ul>
<div class="navbox-styles"><style data-mw-deduplicate="TemplateStyles:r1236075235">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbox{box-sizing:border-box;border:1px solid #a2a9b1;width:100%;clear:both;font-size:88%;text-align:center;padding:1px;margin:1em auto 0}.mw-parser-output .navbox .navbox{margin-top:0}.mw-parser-output .navbox+.navbox,.mw-parser-output .navbox+.navbox-styles+.navbox{margin-top:-1px}.mw-parser-output .navbox-inner,.mw-parser-output .navbox-subgroup{width:100%}.mw-parser-output .navbox-group,.mw-parser-output .navbox-title,.mw-parser-output .navbox-abovebelow{padding:0.25em 1em;line-height:1.5em;text-align:center}.mw-parser-output .navbox-group{white-space:nowrap;text-align:right}.mw-parser-output .navbox,.mw-parser-output .navbox-subgroup{background-color:#fdfdfd}.mw-parser-output .navbox-list{line-height:1.5em;border-color:#fdfdfd}.mw-parser-output .navbox-list-with-group{text-align:left;border-left-width:2px;border-left-style:solid}.mw-parser-output tr+tr>.navbox-abovebelow,.mw-parser-output tr+tr>.navbox-group,.mw-parser-output tr+tr>.navbox-image,.mw-parser-output tr+tr>.navbox-list{border-top:2px solid #fdfdfd}.mw-parser-output .navbox-title{background-color:#ccf}.mw-parser-output .navbox-abovebelow,.mw-parser-output .navbox-group,.mw-parser-output .navbox-subgroup .navbox-title{background-color:#ddf}.mw-parser-output .navbox-subgroup .navbox-group,.mw-parser-output .navbox-subgroup .navbox-abovebelow{background-color:#e6e6ff}.mw-parser-output .navbox-even{background-color:#f7f7f7}.mw-parser-output .navbox-odd{background-color:transparent}.mw-parser-output .navbox .hlist td dl,.mw-parser-output .navbox .hlist td ol,.mw-parser-output .navbox .hlist td ul,.mw-parser-output .navbox td.hlist dl,.mw-parser-output .navbox td.hlist ol,.mw-parser-output .navbox td.hlist ul{padding:0.125em 0}.mw-parser-output .navbox .navbar{display:block;font-size:100%}.mw-parser-output .navbox-title .navbar{float:left;text-align:left;margin-right:0.5em}body.skin--responsive .mw-parser-output .navbox-image img{max-width:none!important}@media print{body.ns-0 .mw-parser-output .navbox{display:none!important}}


/* end https://en.wikipedia.org/ */
</style></div><div role="navigation" class="navbox" aria-labelledby="Existential_risk_from_artificial_intelligence277" style="padding:3px"><table class="nowraplinks mw-collapsible expanded navbox-inner" style="border-spacing:0;background:transparent;color:inherit"><tbody><tr><th scope="col" class="navbox-title" colspan="2"><div id="Existential_risk_from_artificial_intelligence277" style="font-size:114%;margin:0 4em"><a href="Existential_risk_from_artificial_intelligence" title="Existential risk from artificial intelligence">Existential risk</a> from <a href="Artificial_intelligence" title="Artificial intelligence">artificial intelligence</a></div></th></tr><tr><th scope="row" class="navbox-group" style="width:1%">Concepts</th><td class="navbox-list-with-group navbox-list navbox-odd hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Artificial_general_intelligence" title="Artificial general intelligence">AGI</a></li>

<li><a href="AI_boom" title="AI boom">AI boom</a></li>
<li><a href="AI_capability_control" title="AI capability control">AI capability control</a></li>
<li><a href="AI_safety" title="AI safety">AI safety</a></li>
<li><a href="AI_takeover" title="AI takeover">AI takeover</a></li>
<li><a href="Consequentialism" title="Consequentialism">Consequentialism</a></li>
<li><a href="Effective_accelerationism" title="Effective accelerationism">Effective accelerationism</a></li>
<li><a href="Ethics_of_artificial_intelligence" title="Ethics of artificial intelligence">Ethics of artificial intelligence</a></li>
<li><a href="Existential_risk_from_artificial_intelligence" title="Existential risk from artificial intelligence">Existential risk from artificial intelligence</a></li>
<li><a href="Friendly_artificial_intelligence" title="Friendly artificial intelligence">Friendly artificial intelligence</a></li>
<li><a href="Instrumental_convergence" title="Instrumental convergence">Instrumental convergence</a></li>
<li><a href="Vulnerable_world_hypothesis" title="Vulnerable world hypothesis">Vulnerable world hypothesis</a></li>
<li><a href="Intelligence_explosion" class="mw-redirect" title="Intelligence explosion">Intelligence explosion</a></li>
<li><a href="Longtermism" title="Longtermism">Longtermism</a></li>
<li><a href="Machine_ethics" title="Machine ethics">Machine ethics</a></li>
<li><a href="Risk_of_astronomical_suffering" title="Risk of astronomical suffering">Suffering risks</a></li>
<li><a href="Superintelligence" title="Superintelligence">Superintelligence</a></li>
<li><a href="Technological_singularity" title="Technological singularity">Technological singularity</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">Organizations</th><td class="navbox-list-with-group navbox-list navbox-even hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Alignment_Research_Center" title="Alignment Research Center">Alignment Research Center</a></li>
<li><a href="Center_for_AI_Safety" title="Center for AI Safety">Center for AI Safety</a></li>
<li><a href="Center_for_Applied_Rationality" title="Center for Applied Rationality">Center for Applied Rationality</a></li>
<li><a href="Center_for_Human-Compatible_Artificial_Intelligence" title="Center for Human-Compatible Artificial Intelligence">Center for Human-Compatible Artificial Intelligence</a></li>
<li><a href="Centre_for_the_Study_of_Existential_Risk" title="Centre for the Study of Existential Risk">Centre for the Study of Existential Risk</a></li>
<li><a href="EleutherAI" title="EleutherAI">EleutherAI</a></li>
<li><a href="Future_of_Humanity_Institute" title="Future of Humanity Institute">Future of Humanity Institute</a></li>
<li><a href="Future_of_Life_Institute" title="Future of Life Institute">Future of Life Institute</a></li>
<li><a href="Google_DeepMind" title="Google DeepMind">Google DeepMind</a></li>
<li><a href="Humanity%2B" title="Humanity+">Humanity+</a></li>
<li><a href="Institute_for_Ethics_and_Emerging_Technologies" title="Institute for Ethics and Emerging Technologies">Institute for Ethics and Emerging Technologies</a></li>
<li><a href="Leverhulme_Centre_for_the_Future_of_Intelligence" title="Leverhulme Centre for the Future of Intelligence">Leverhulme Centre for the Future of Intelligence</a></li>
<li><a href="Machine_Intelligence_Research_Institute" title="Machine Intelligence Research Institute">Machine Intelligence Research Institute</a></li>
<li><a href="OpenAI" title="OpenAI">OpenAI</a></li>
<li><a href="Safe_Superintelligence_Inc." title="Safe Superintelligence Inc.">Safe Superintelligence</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">People</th><td class="navbox-list-with-group navbox-list navbox-odd hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Slate_Star_Codex" title="Slate Star Codex">Scott Alexander</a></li>
<li><a href="Sam_Altman" title="Sam Altman">Sam Altman</a></li>
<li><a href="Yoshua_Bengio" title="Yoshua Bengio">Yoshua Bengio</a></li>
<li><a href="Nick_Bostrom" title="Nick Bostrom">Nick Bostrom</a></li>
<li><a href="Paul_Christiano_(researcher)" class="mw-redirect" title="Paul Christiano (researcher)">Paul Christiano</a></li>
<li><a href="K._Eric_Drexler" title="K. Eric Drexler">Eric Drexler</a></li>
<li><a href="Sam_Harris" title="Sam Harris">Sam Harris</a></li>
<li><a href="Stephen_Hawking" title="Stephen Hawking">Stephen Hawking</a></li>
<li><a href="Dan_Hendrycks" title="Dan Hendrycks">Dan Hendrycks</a></li>
<li><a href="Geoffrey_Hinton" title="Geoffrey Hinton">Geoffrey Hinton</a></li>
<li><a href="Bill_Joy" title="Bill Joy">Bill Joy</a></li>
<li><a href="Shane_Legg" title="Shane Legg">Shane Legg</a></li>
<li><a href="Elon_Musk" title="Elon Musk">Elon Musk</a></li>
<li><a href="Steve_Omohundro" title="Steve Omohundro">Steve Omohundro</a></li>
<li><a href="Huw_Price" title="Huw Price">Huw Price</a></li>
<li><a href="Martin_Rees" title="Martin Rees">Martin Rees</a></li>
<li><a href="Stuart_J._Russell" title="Stuart J. Russell">Stuart J. Russell</a></li>
<li><a href="Ilya_Sutskever" title="Ilya Sutskever">Ilya Sutskever</a></li>
<li><a href="Jaan_Tallinn" title="Jaan Tallinn">Jaan Tallinn</a></li>
<li><a href="Max_Tegmark" title="Max Tegmark">Max Tegmark</a></li>
<li><a href="Frank_Wilczek" title="Frank Wilczek">Frank Wilczek</a></li>
<li><a href="Roman_Yampolskiy" title="Roman Yampolskiy">Roman Yampolskiy</a></li>
<li><a href="Eliezer_Yudkowsky" title="Eliezer Yudkowsky">Eliezer Yudkowsky</a></li></ul>
</div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">Other</th><td class="navbox-list-with-group navbox-list navbox-even hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Artificial_Intelligence_Act" title="Artificial Intelligence Act">Artificial Intelligence Act</a></li>
<li><i><a href="Do_You_Trust_This_Computer%3F" title="Do You Trust This Computer?">Do You Trust This Computer?</a></i></li>
<li><i><a href="Human_Compatible" title="Human Compatible">Human Compatible</a></i></li>
<li><a href="Open_letter_on_artificial_intelligence_(2015)" class="mw-redirect" title="Open letter on artificial intelligence (2015)">Open letter on artificial intelligence (2015)</a></li>
<li><i><a href="Our_Final_Invention" title="Our Final Invention">Our Final Invention</a></i></li>
<li><a href="Roko's_basilisk" title="Roko's basilisk">Roko's basilisk</a></li>
<li><a href="Statement_on_AI_risk_of_extinction" class="mw-redirect" title="Statement on AI risk of extinction">Statement on AI risk of extinction</a></li>
<li><i><a href="Superintelligence%3A_Paths%2C_Dangers%2C_Strategies" title="Superintelligence: Paths, Dangers, Strategies">Superintelligence: Paths, Dangers, Strategies</a></i></li>
<li><i><a href="The_Precipice%3A_Existential_Risk_and_the_Future_of_Humanity" title="The Precipice: Existential Risk and the Future of Humanity">The Precipice</a></i></li></ul>
</div></td></tr><tr><td class="navbox-abovebelow" colspan="2"><div><span class="noviewer" typeof="mw:File"><span title="Category"></span></span> Category</div></td></tr></tbody></table></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-07-21" href="https://en.wikipedia.org/wiki/?title=AI_alignment&amp;oldid=1301770021">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>

</body></html>